Compare commits

..
Author SHA1 Message Date
jschoubben 2b5119ecd2 Issue 114 is answered: the controller is a process, by ADR 0142
It asked a question rather than reporting a defect, and the question was taken
five days later — the mesh's own components are binaries on the machine, and
third-party software stays a container because an image is the right way to
carry somebody else's build. The delivery of them is issue 142.
2026-09-29 22:38:11 +02:00
jschoubben 96bdffa9bc Two records were numbered 127; the second becomes 149, and is resolved
A number identifies a record, and two were given 127 on 2026-09-27. The one
three documents and three source files cite by number keeps it; the other
becomes 149, says so in its own heading, and its two inbound references are
repointed.

It is also resolved: an empty declaration is sent carrying owns_nothing rather
than skipped, and the host refuses an empty body that does not carry it, so
emptiness cannot be read as a truncated declaration.
2026-09-29 22:33:40 +02:00
jschoubben 4eb16f1028 Three proposed records were already built; two are still yours to call
0037 (where a module lives), 0113 (the vault makes every secret) and 0115 (one
assignment of a module per node) described arrangements the mesh has — a
catalogue of other people's software plus a table filled by module add, a vault
answering a secret requirement six modules make, and a rule the assignment
table's primary key already enforces. Each is accepted against what was built,
and says so in its own words.

0037's other half is not built: a manifest outside this catalogue has no check,
which is issue 148.

0068 (the lab takes requests) and 0114 (a shared credential rotates over two
credentials) stay proposed. Neither is built, and both are decisions rather
than records of something that happened.
2026-09-29 22:31:55 +02:00
jschoubben d199de40db Merge the 146/147 records, which 006's note links to 2026-09-29 22:31:46 +02:00
jschoubben 9a1dc4665c Grooming: issue 006's knowledge base is the predecessor's, and is gone
The record catches README claiming this repository is indexed into a knowledge
base. Nothing indexed it, and since the cut-over there is nothing to index it
into — the surface that answered is on the transport the mesh removed (issue
147). Noted where the record is, so the next reader does not go looking for a
search that cannot exist.
2026-09-29 22:26:13 +02:00
jschoubben 14be8576f8 Grooming: five issues were fixed and never closed, and one is not
088, 089, 120, 128 and 130 each name a commit that is on main and cites them —
the forge's address following a moved port, a route naming its endpoint, a
provisioner asking the backend what is there, the hosts file written into a
marked block, and undeclaring giving a unit back the state it was found in.
Each says it was closed by reading commits rather than by a run, so nobody
reads a green that was never measured.

129 stays located on purpose: ca-trust is merged and no machine holds it, so
the symptom it opened on is still true everywhere.
2026-09-29 22:25:25 +02:00
jschoubben 3c0f7082e6 Issue 147: the tool surface is not the mesh's, it is the predecessor's
Diagnosed from the configuration: the tool server is HAL's brain, a local
process on the workstation with the predecessor's broker URL in the assistant's
own config. No manifest, no assignment, no seat, no account. Nothing regressed
— the mesh removed a transport this program still dials, and the program was
never part of the mesh. The mesh has a tool model and nothing publishes an
operator-facing surface onto it.
2026-09-29 21:50:50 +02:00
jschoubben 0dd00e88b6 Issue 147: the operator's tools still dial the bus that was removed
Every tool call fails with 'AMQP not connected', on every node including the
local one, because the tool surface still opens an AMQP connection and that
transport was deleted at the cut-over. The mesh reports healthy throughout —
what broke is the thing standing outside asking it questions, so nothing the
mesh checks is about it.
2026-09-29 21:39:55 +02:00
jschoubben eef54917ec Issue 146: the double enrolment was two consumers sharing a delivery subject
Not about enrolment. A push consumer delivers onto an ordinary subject and
everything subscribed to it gets a copy; the controller's two consumers were
both named after it, so both were given the same subject and the one process
acted on every message twice. Enrolment is where it drew blood because a second
enrolment mints a second credential.
2026-09-29 21:32:52 +02:00
jschoubben e9b1010bc0 Issue 146: what made it slow, and what was changed so it is not 2026-09-29 17:45:25 +02:00
jschoubben f6ed3545b7 Issue 146: a first node now enrols, and is enrolled twice
Two more faults behind the three already fixed. The account a token is the
password of was never recorded, and the comment above the issuing code said
it was; issuing now records it. Placing the composed list at genesis is the
other half — the control plane says what it composed and whoever raises the
machine writes it beside the bus, because no declaration can reach a machine
that has not enrolled.

With that a first node enrols. It is then enrolled twice from one attempt,
each minting a credential, and it keeps the answer to the first while the mesh
keeps the second. The trail for that one stops at a duplicate that survived
message-id deduplication.
2026-09-29 17:36:40 +02:00
jschoubben 0b08cdfce1 Merge pull request 'ADR 0147: a module anchors the mesh's authority, and issue 146: the foundation cannot be raised' (#184) from decision/0147-a-module-anchors-the-meshs-authority into main 2026-09-29 14:06:41 +00:00
jschoubben bffd2af40c Issue 146 diagnosed: four faults stacked, three fixed, the fourth is genesis
Raising a first node hits them in order: the bundle's bus image named for a
registry that is gone; the certificate made by openssl in an image that has
none; enrolment dialling TLS at a bus that speaks first, then refusing its own
token for the empty bus. Each was right until the bus changed and nothing has
raised a foundation since.

The fourth is not a patch: the installer carries the bus's first user list and
the controller composes the rest through the bus being a module, and at genesis
there is no module — so the first node cannot be let onto the bus it just
raised.
2026-09-29 15:42:32 +02:00
jschoubben 4a51ea4b3a Issue 146: the foundation cannot be raised on the bus the mesh runs on
Found trying to run ADR 0147's bed. The older bundle raises the previous
broker and a control plane that refuses to start without MESH_BUS_NATS; the
newer one stops a step earlier, asking for a certificate from openssl in an
image that has none. No bed can run while this holds, so 0147's check section
now says what actually stands behind it — the rendering, not a machine.
2026-09-29 15:25:58 +02:00
jschoubben 6c14d313b8 ADR 0147: a module anchors the mesh's authority on a machine, and takes it away again
Issue 129: every internal HTTPS name fails verification on every machine,
because nothing has ever written the mesh's root into a trust store. The
report proposed the controller inject it the way the private network writes
the registry's trust; this record rejects that — reachability and trust are
not the same fact, and where anchors live is the host's difference, not the
controller's. A module requiring internal-acme-ca does the whole of it, and
being unassigned undoes it.
2026-09-29 15:02:37 +02:00
mesh-admin ced547dae9 Merge pull request 'ADR 0146: connectivity is checked by name, per hosting form' (#183) from decision/0146-connectivity-by-name into main 2026-09-29 12:01:57 +00:00
jschoubben 38482435af ADR 0146: connectivity is checked by name, per hosting form, with a valid certificate
0145's core stands — a module checks what the mesh claims, from where the callers
are — and everything it said about what to dial was wrong.

A raw port is not how anything in this mesh is reached. A real caller resolves a
name, the proxy answers, the proxy reaches the service; dialling a port tests the
last hop of a four-hop path and skips the three where most connectivity lives.

So each hosting form gets its own endpoint, route and name:
connect-docker.<node>.internal for a container, connect-process for the mesh's own
code in a unit it writes, connect-unit for a unit a package ships — and the same
under each public domain. Every machine checks every machine, by name, over TLS,
verifying the certificate against the authority that should have issued it. The
names are the instrument: a failure reads as connect-docker.g14.internal did not
answer.

No name is written anywhere: the machines come from the roster fact, the labels are
the module's, and a machine that joins is checked by the others on the next push.

Five things this needs that the mesh does not have, named rather than assumed —
a module every machine has (ScopeNode is exclusivity, not obligation), a container
running a bundle, a unit a package ships, the public domain in the roster fact, and
issue 129, until which every internal name will correctly fail verification.
2026-09-29 14:01:55 +02:00
mesh-admin cf8a8d78c9 Merge pull request 'ADR 0145: a module checks what the mesh claims is reachable' (#182) from decision/0145-a-module-checks-what-the-mesh-claims into main 2026-09-29 11:44:58 +00:00
jschoubben 9de25994e9 ADR 0145: a module checks what the mesh claims is reachable, and it checks itself
The mesh asserts three things are callable and has never checked any of them. A
module on every machine serves an endpoint of its own and dials every other
machine's, from the position the callers are in — not the host and not the control
plane, both of which reach those addresses by paths no ordinary caller uses and
would have passed throughout the outage.

Its probe is its own endpoint, declared reachable over the private network, so it
is admitted by exactly the rule that governs every internally-exposed service. The
tempting target is a service every machine has, and those are the ones never closed
— ssh above all — which would have passed for all eleven hours.

Resolution and connection reported separately, because they have different owners.
One failure is not a fault, and the count travels with the result. It reports and
repairs nothing.

Built on a scheduled container, a rendered roster fact and a bundle: no new
vocabulary. Adopted on its own merits, which is what 0143 was not.
2026-09-29 13:44:56 +02:00
mesh-admin 64ea47b11d Merge pull request 'ADR 0144: anything on a machine may call anything on it' (#181) from decision/0144-local-is-not-a-boundary into main 2026-09-29 11:32:03 +00:00
jschoubben eba24a72af ADR 0144: anything on a machine may call anything on it, superseding 0143
Everything should be able to call what runs on the same machine, another machine's
service exposed to the private network, and another machine's service exposed
publicly. Three cases; the filter had two.

The first was expressed as the machines' own addresses on the private network. A
caller on the machine carries such an address; a caller in one of its containers
carries a bridge address and matched nothing — measured, same destination and same
machine: src 10.10.0.1 against src 172.17.0.8. The second case worked by accident,
because the tunnel rewrites a caller's address to the sending machine's. Two of
three working is why it read as correct.

0143 answered the wrong question. It proposed verifying each grant from the
consumer's own network position and went to length about which position, because
whether a caller sat in a container changed the answer — and that difference was the
bug. Observing a configuration error is not its remedy. Superseded, and nothing
replaces it; whether the mesh should check a grant is still open in issue 145 and
must stand on its own.

And a module is not a container: 61 of 72 happen to use one, 11 do not, and a rule
reasoning about containers describes most of the mesh rather than the mesh.
2026-09-29 13:32:01 +02:00
mesh-admin 78d4873f4f Merge pull request 'ADR 0143: a consumer verifies the grant it is given' (#180) from decision/0143-a-consumer-verifies-its-grant into main 2026-09-29 11:07:20 +00:00
jschoubben ddfd62edf6 ADR 0143: a consumer verifies the grant it is given
A grant is four facts and a credential — the provision, the machine, the port, and
who the consumer is when it connects — and it is the whole mechanism by which
anything in the mesh reaches anything else. The mesh asserted it and never found
out whether it was true. Issue 145 is what that cost: eleven hours of 'every module
current with its source' while a module could not reach its database.

The consumer verifies it, from its own network position. Not the control plane,
which reaches the address by a path no consumer uses. Not the machine, whose own
packets carried a source address the filter admitted while every container's was
refused — so a host-side check would have passed throughout the outage it exists to
catch. That is inference from the rule that was loaded, stated as such.

A connection and nothing more; speaking each provision's protocol would be a second
implementation of every provision. One failure is not news, only consecutive ones,
and the count is reported rather than the last attempt. A consumer that is not
running reads unchecked, which is a different sentence from broken.

What it costs to be wrong is the constraint on all of it: a check that calls a
working provision broken trains a reader to ignore the report.
2026-09-29 13:07:18 +02:00
mesh-admin f65664640a Merge pull request 'Issue 145: a machine reads healthy while its modules cannot reach each other' (#179) from issue/145-healthy-while-broken into main 2026-09-29 10:37:10 +00:00
jschoubben 492ac7be18 Issue 145: a machine reads healthy while its modules cannot reach each other
Converging the control node closed every path by which a module reached another by
the machine's own name, and it ran for eleven hours while the mesh answered 'all
doing what they were told, all heard from, running what the mesh would send them,
and every module current with its source'.

6,154 database failures in one module's log, beginning at the minute of the flip.
The service accepted TCP and never answered HTTP. Every check the mesh makes
passed, because every check the mesh makes is about the relationship between the
mesh and a machine — applied, current, containers running. None asks whether a
module can reach what it requires, though the mesh composes every grant and so
knows exactly who requires what.

The filter fault is fixed. The eleven hours are the measurement, not the bug.

Also records on 144 that 'more closed than the mesh believes' was harmless only
for what is reached from outside, and an outage for what is reached from within.
2026-09-29 12:37:08 +02:00
mesh-admin 5ac77e3cef Merge pull request 'ADR 0138: reach asks for names on a routed endpoint' (#178) from decision/0138-insight-reach-and-the-proxy into main 2026-09-29 00:50:28 +00:00
jschoubben a619022c35 ADR 0138: a progressive insight — reach asks for names on a routed endpoint
The record says internal and public each mean something to the filter. For an
endpoint the proxy serves, the second half is wrong, and ADR 0045 said so first:
a public service is exposed through the proxy, listening from the mesh, not by
opening its own port.

Found by trying to express one real module, not by review — routed name public
because browsers post to it, machine port private because it serves a dashboard in
cleartext. Under one value for both, saying public would have reopened a port an
operator had just closed. Measured the same evening: the routed name answered from
the internet over TLS while the port was refused from the same place.

Corrects a fact. One statement per endpoint with three things derived from it
stands; the filter column applies to an unrouted endpoint.
2026-09-29 02:50:25 +02:00
mesh-admin 49c065c204 Merge pull request 'Issue 143: correct the diagnosis' (#177) from issue/143-corrected-diagnosis into main 2026-09-29 00:33:13 +00:00
jschoubben ddb3980f09 Issue 143: correct the diagnosis — the step exists and did not fire
The first account said the declaration carries no resource that would disable the
found firewall and that nothing implemented the sentence the flip prints. Wrong:
the mechanism is a step in the host's own apply, retireFirewall, and it is careful
— it reads back that the mesh's table is loaded before retiring anything, records
the forward policies first, and verifies ufw reports inactive afterwards.

What is established: ufw was active and enabled two minutes after a flip that
reported the node converged with nothing failed; and the machine's record now
reads disabled_by_mesh: true, written by a reconcile fifty minutes later that
found ufw already inactive because an operator had disabled it by hand.

So the step did not take effect and the record says it did. The candidates are
named rather than chosen — the step is skipped silently when the apply has any
failure, the mesh's table is loaded by a service in the same apply so ordering is
open, and the host's detail lines do not reach the journal, which is why this has
candidates instead of a cause.
2026-09-29 02:33:10 +02:00
mesh-admin 75a8f6abc7 Merge pull request 'Issues 143 and 144: the found firewall is neither retired nor all of it' (#176) from issue/143-and-144-the-found-firewall into main 2026-09-28 23:36:10 +00:00
jschoubben 862253518f Issues 143 and 144: the found firewall is neither retired nor all of it
Both found by converging the control node — the first machine with a firewall to
flip, since the two before it had none.

143: the preview and the flip both say the found firewall is disabled. The node
reported converged, 372 resources applied, nothing failed, and ufw is still
enabled and active. A converged node's declaration carries no resource that would
disable it; the sentence is printed by the command and nothing implements it.

144: ufw was never what filtered the traffic that mattered there. Fifty forwarded
openings converged through it had matched zero packets, while a chain the
predecessor installed in the container runtime's pre-accept hook did the work —
in memory only, recreated by nothing. The mesh's filter now covers that path, so
the machine no longer depends on it, but the chain remains and is the only thing
refusing the bus and the registry, which the design requires reachable from
anywhere so a machine can enrol before it has a private address.
2026-09-29 01:36:08 +02:00
mesh-admin 51ef3eb7e2 Merge pull request 'ADR 0142: the mesh delivers its own components as binaries' (#175) from decision/0142-mesh-delivers-its-own-components into main 2026-09-28 22:51:44 +00:00
jschoubben 346e613995 ADR 0142: the mesh delivers its own components as binaries, not container images
Measured: the host is a binary somebody copied to four machines, owned by no
package and built by nothing, while the controller, catalogue, builder and vault
are container images publishing no ports at all. Same language, same project,
same kind of work, delivered two ways — and the difference is not a judgement
about either, it is that images are the only delivery that works.

What it costs: genesis must raise a container runtime before the control plane
can exist; updating the control plane goes through a registry the control plane
runs; a host change cannot be rolled out at all; and compiling the language the
mesh is written in is not a capability of the builder, so the controller is built
from a hand-written Dockerfile — the incantation the bundle toolchain exists to
abolish.

Third-party software stays a container: the store, the registry, the broker are
somebody else's build. The container runtime stays on the machine for modules.
What changes is that the control plane no longer needs it to exist.

The receiving half is already built and tested (ADR 0141). Staged: compile Go, an
artifact names its target, deliver a binary, the host first, then the rest, genesis
last.
2026-09-29 00:51:42 +02:00
mesh-admin 25ca9898d5 Merge pull request 'ADR 0141: a progressive insight on what delivery costs' (#174) from decision/0141-progressive-insight-on-delivery into main 2026-09-28 22:32:17 +00:00
jschoubben 709c095387 ADR 0141: a progressive insight — the delivery is not 'nothing new'
The record claimed a version reaches a machine as an ordinary archive with
nothing new needed. Two things it needs do not exist: no toolchain can compile
the host (the list is typescript and python, and the control plane, also Go, is
built as an image from a Dockerfile instead), and nothing interpolates a built
version into a resource path, so nothing can ask for .../versions/<version>/.

The decision, the options weighed and every consequence stand — the host half is
merged and tested. What was understated was the cost, so it is corrected in place
and dated rather than superseded.
2026-09-29 00:32:14 +02:00
mesh-admin c262de3833 Merge pull request 'ADR 0141: the host delivers its own successor' (#173) from decision/0141-the-host-delivers-its-own-successor into main 2026-09-28 22:22:28 +00:00
jschoubben a815433214 ADR 0141: the host delivers its own successor, and versions live side by side
The supervision was already right — a clean exit means the host stood aside and
the launcher runs what is on disk, failures are counted, and a rollback happens
at the limit. Two things made it dead code: nothing told the running host a
successor was waiting, and the rollback resolved its known-good version through
pacman, which no machine here uses and which two of three operating systems do
not have.

Keeping a version rather than a path was the clue. Versions live side by side in
directories named for them; the newest runs; the running one stands aside between
reconciles; a reconcile that completes records itself and retires what is older
than its predecessor; rollback starts that predecessor. No new resource kind and
nothing new on the bus — an archive already fetches by digest, and the path
written is never the path executing.

Answers issue 142.
2026-09-29 00:22:04 +02:00
mesh-admin 61dc90e508 Merge pull request 'Issue 142: the host is the one thing the mesh does not deliver' (#172) from issue/142-the-host-is-not-delivered into main 2026-09-28 22:03:58 +00:00
jschoubben a7cf5c0c1b Issue 142: the host is the one thing the mesh does not deliver
A host change merged yesterday reached no machine without a person copying a
file. The host is not a build target, no declaration delivers it, and the half
that recovers from a bad host — noticing the executable changed, a known-good
record, a launcher that rolls back — is written, tested and called by nothing.
All four machines run a byte-identical hand-copied binary that no package owns
and no record names, so nothing can say a machine is behind.

Found because ADR 0140 needs the machine to report a new fact, and merging that
could not roll it out.
2026-09-29 00:03:41 +02:00
mesh-admin c6f86ae935 Merge pull request 'ADR 0140: the filter constrains what arrives from outside' (#171) from decision/0140-filter-constrains-what-arrives-from-outside into main 2026-09-28 21:32:10 +00:00
jschoubben 1aeb4fe8d8 ADR 0140: the filter constrains what arrives from outside, and says nothing about a machine's own guests
Reading a converged machine's rendered rules showed the cause: the chain blocks
everything passing through and then allows the machine's own containers back by
listing their address ranges. 0137 made that list typeable and 0139 tried to
generate it; both refined a list that should not exist, because the mesh has no
position on a container reaching outward. Constrain what arrives from outside,
allow what did not, and let the machine report which links face outside — one
fact instead of a list. Ports keep following the modules unchanged.

The records check now allows one record to supersede several, and stops
requiring a withdrawn record's own citations to be live.
2026-09-28 23:31:50 +02:00
mesh-admin 08108b569d Merge pull request 'ADRs 0138 and 0139: an endpoint's reach, and networks forwarded because a module declared them' (#170) from decision/0138-endpoint-reach-and-0139-declared-networks into main 2026-09-28 21:03:49 +00:00
jschoubben 14ff89fa40 ADRs 0138 and 0139: an assignment binds an endpoint and says how far it reaches, and a network is forwarded because a module declared it
Both follow from the same rule the mesh is built on — a node's configuration is
composed from the modules assigned to it. Reach was settled separately by the
filter, the proxy's names and the certificate authority, so "this must not be
public" could not be written; it becomes one value on the assignment that all
three read. And the forward chain consulted two constants plus a typed list
although modules already declare their networks; it now forwards what they
declared, with the host rendering the addresses it allocated.
2026-09-28 23:03:32 +02:00
mesh-admin 2126e7b2cb Merge pull request 'Issues 140 and 141: an endpoint's reach, and a forward chain that does not follow the modules' (#169) from issue/140-endpoint-reach-and-141-forward-chain into main 2026-09-28 20:58:03 +00:00
jschoubben dcdfcf104e Issues 140 and 141: an endpoint's reach is declared nowhere, and the forward chain follows constants instead of the modules
Found preparing the control-node's convergence. Reach is settled independently by
the filter, the proxy's names and the certificate authority, so "this must not be
public" cannot be written and a public certificate is obtained regardless. And the
forward chain allows two hardcoded ranges plus a typed list, though the mesh
already knows which networks exist because its own modules declared them — a range
wide enough to keep four of them would have forwarded two predecessor leftovers too.
2026-09-28 22:57:38 +02:00
mesh-admin d23ace1646 Merge pull request 'Issues 138 and 139' (#168) from issue/138-the-uplink-seat-and-139-an-internal-route-name into main 2026-09-28 19:46:22 +00:00
jschoubben 5dbde0b13a Issues 138 and 139: a seat with interchangeable holders that are not, and an internal route name that resolves to the wrong machine 2026-09-28 21:46:20 +02:00
mesh-admin db3868e2b0 Merge pull request 'ADR 0137: a machine says which networks it routes' (#167) from decision/0137-a-machine-says-which-networks-it-routes into main 2026-09-28 19:31:03 +00:00
jschoubben 05039c4f10 ADR 0137: a machine says which networks it routes, and issue 137 is how that was found 2026-09-28 21:31:01 +02:00
mesh-admin b7bf601ca4 Merge pull request 'Issue 136: a module may name a program the machine does not have' (#166) from issue/136-a-module-may-name-a-program-the-machine-lacks into main 2026-09-28 18:50:37 +00:00
jschoubben bff32e3370 Issue 136: a module may name a program the machine does not have, and everything reports success 2026-09-28 20:50:35 +02:00
mesh-admin d6b62387f2 Merge pull request 'Issue 135: a container's mesh names are not compared' (#165) from issue/135-a-containers-mesh-names into main 2026-09-28 15:19:37 +00:00
jschoubben 4789624857 Issue 135: a container's mesh names are not compared, so a moved address is never noticed
One container restarted 2286 times over five days while the mesh reported the machine as doing what it
was told. Its overlay address was five days out of date: the host compares a container by a digest of
its spec, and the mesh's names were not in it, so a container whose image and files never changed was
left alone holding a name that no longer resolved. Forty-eight others were current only because
something else had recreated them.

The same fault as issue 045, in the field that was left out. Resolved by putting the names in the
digest.
2026-09-28 17:19:35 +02:00
mesh-admin e00862e317 Merge pull request 'ADR 0112 is accepted, and issue 134 records what it is not yet' (#164) from decision/0112-accepted-and-issue-134 into main 2026-09-28 15:12:15 +00:00
jschoubben 3a92e4b80c ADR 0112 is accepted, and issue 134 records what it is not yet
0120 was already accepted; what failed the check was that it rests on 0112, still marked proposed —
and so do four designs. The decision stands: a definition names no node, no mesh and no host path, and
everything a module needs is a requirement the mesh resolves.

Accepting it makes the gap visible rather than hiding it, so issue 134 states it. 0112 says how it is
checked — 'a catalogue test finds no domain name in any definition value' — and there is no such test.
Asked by hand: seven modules name this installation in a value the mesh acts on, and eight mention a
public name in prose nothing reads. The two are not the same fault and the fixes differ, which is why
the issue separates them rather than counting to fifteen.
2026-09-28 17:12:12 +02:00
mesh-admin bfddf78bf3 Merge pull request 'Design 32: what shipped today, and the records it rests on' (#163) from design/28-and-32-what-shipped into main 2026-09-28 14:50:44 +00:00
jschoubben 0b291d89f3 Design 32: what shipped today, and the records it rests on
Two of its statements are built — a version preparing its state, and the mesh saying what it applied —
so the document is in-progress rather than proposed, and names the code that owns them. Resting on a
live record rather than a superseded one: 0127 was replaced by 0131.

What this exposes is pre-existing: it also rests on ADR 0112, which is still proposed, and a document
that is not itself proposed may not. ADR 0120 has rested on it the same way for a while. Accepting or
superseding 0112 is a decision, not a cleanup, so it stays visible in the check rather than papered
over.
2026-09-28 16:50:41 +02:00
mesh-admin a924efcc28 Merge pull request 'ADR 0136: a step gates its module, not the machine' (#162) from decision/0136-a-step-gates-its-module into main 2026-09-28 13:38:18 +00:00
jschoubben b421c73a5a ADR 0136: a step gates its module, not the machine
ADR 0135 made a step something the mesh derives for any module that prepares its state, which turned
ADR 0052's reach into a fault: a module whose database is briefly unreachable would stop every module
declared after it on that machine — the fault issue 011 already removed for every other shape, and
the reason the catalogue migrates itself at start rather than in a step.

A step now stops the rest of its own module and nothing else; an action still gates the machine,
because genesis is a row of them and they belong to no module. What was not attempted is reported as
skipped rather than left to be inferred from silence.
2026-09-28 15:38:16 +02:00
mesh-admin 017b1401e4 Merge pull request 'ADR 0135: a module version prepares its state before it runs' (#161) from decision/0133-0134-migrations-and-deploy-facts into main 2026-09-28 09:57:47 +00:00
jschoubben 476cda417d ADR 0135 supersedes 0133: a module version prepares its state before it runs
Two faults in 0133, both caught on review. It put the declaration on a container — one resource kind
the host applies — so every author would restate the machine's arrangement and a module's own
lifecycle would be tied to how its artifact happens to run. A module declares entrypoints for its
tools and its provisioner; preparing its state is the same vocabulary and nothing about a runtime.

And it derived the scope from the machine, which the facts already answer: a consumer is a module on
a machine (issue 022, migration 0015), so what the mesh provisions is per consumer. A module on three
machines has three databases, there is no shared state to race over, and the lock obligation 0133
invented was for a situation the mesh does not produce. The level question HAL answered with stages
dissolves — the scope of preparation is the scope of the state, and the mesh knows it.

0133 keeps its reasoning and gains a pointer; design 32 and issue 133 name the live record.
2026-09-28 11:57:16 +02:00
mesh-admin 74ba3ff1d4 Merge pull request 'ADR 0133 and 0134: who runs migrations, and the mesh saying what it applied' (#160) from decision/0133-0134-migrations-and-deploy-facts into main 2026-09-28 09:45:49 +00:00
jschoubben 891c7a945e ADR 0133 and 0134: who runs migrations, and the mesh saying what it applied
0133 — a module owns its migrations and the mesh owns when they run. A container declares what must
run before it; the mesh derives the gated step from the resource it precedes, so the image, the
environment and the credentials come from the one place they are described. The module owns the SQL,
the dialect and the lock; the mesh owns the moment and refuses to start a version whose step failed.
Per node, with no level: a step that ran once somewhere leaves every other machine ungated, and
'once, mesh-wide' is what holding a seat already means.

0134 — the pipeline is observable from a merge to an artifact and goes dark at the machine. What a
node now runs, and what it refused, become facts under the control plane's own seat, emitted when
what a machine runs changes rather than on every convergence pass.

Design 32's lifecycle carries both; issue 133 points at them as what ends the matter it opened.
2026-09-28 11:45:47 +02:00
mesh-admin 83cbeb8db8 Merge pull request 'Issue 133: the control plane's schema is migrated at birth and never again' (#159) from issue/133-the-control-plane-migrates-before-it-serves into main 2026-09-28 08:27:32 +00:00
jschoubben ab6db9369b Issue 133: the control plane's schema is migrated at birth and never again
The mesh replaced its own control plane with a build carrying a migration, applied none of it, and
then recorded no build for three quarters of an hour while saying everything was fine. ADR 0052
already prescribes the shape — a run-once step that gates the server — and the control plane was the
one module that did not use it.
2026-09-28 10:27:30 +02:00
mesh-admin 84571b4825 Merge pull request 'ADR 0132: a seat carries the tools its holder must serve' (#158) from decision/0132-a-seat-carries-the-tools-its-holder-must-serve into main 2026-09-28 08:17:01 +00:00
jschoubben d57196102d ADR 0132: a seat carries the tools its holder must serve
A role's tools belong to the role, not to whichever module holds it today: the seat declares them
with their schemas, serving them is a condition of occupying the seat, and what the mesh can do
becomes a read of its own records rather than a question nothing answers. A module keeps its own
tools — the same module may run without the seat, and then only its own name is true.

Design 33 follows: the three families, addressing a node-scoped seat, discovery, and what serves
this to an agent.
2026-09-28 10:16:59 +02:00
mesh-admin e6402cf777 Merge pull request 'Issue 132: a module can be recorded without the directory it lives in' (#157) from issue/132-a-module-can-be-recorded-without-its-directory into main 2026-09-28 07:20:07 +00:00
57 changed files with 4272 additions and 40 deletions
+16 -3
View File
@@ -164,6 +164,12 @@ def check_rests_on(failures, records):
# decision is exactly what as-is is for."
if rel(path).startswith("03-DESIGN/00-as-is/"):
continue
# A withdrawn record's citations are history. It instructs nobody -- every reader
# is sent to its superseder -- so what it was built on may itself be withdrawn.
# Refusing that would mean rewriting the lineage of a record whose reasoning is
# the thing the immutability rule protects.
if frontmatter(read(path)).get("status") == "superseded":
continue
# An extension that supersedes legitimately names what it replaced.
this = ADR_FILE.match(os.path.basename(path))
supersedes = records[number]["front"].get("superseded-by", "")
@@ -241,13 +247,20 @@ def check_supersession_symmetry(failures, records):
failures.add("supersession", rel(record["path"]), f"superseder does not exist: {by}")
continue
other = records[match.group(1)]
claims = os.path.basename(str(other["front"].get("supersedes", "")))
if claims != record["name"]:
# `supersedes:` may name one record or several. One decision replacing two is a real
# situation -- two records that built and refined the same wrong mechanism are withdrawn
# by the one record that removes it -- and a check that allows only one would force
# either a chain of pro-forma records or an unmarked supersession.
claimed = other["front"].get("supersedes", "")
if isinstance(claimed, str):
claimed = [claimed] if claimed else []
claims = [os.path.basename(str(entry)) for entry in claimed]
if record["name"] not in claims:
failures.add(
"supersession",
rel(other["path"]),
f"ADR {number} says this supersedes it; this record does not say so "
f"(supersedes: {claims or 'absent'})",
f"(supersedes: {', '.join(claims) or 'absent'})",
)
+17 -1
View File
@@ -1,6 +1,6 @@
---
topic: building it
status: proposed
status: accepted
date: 2026-09-01
deciders: jochen
reconstructed: false
@@ -101,3 +101,19 @@ the digest down after building.
**Whether kind 4 deserves a module at all.** Thirty-five descriptions that say *install this and
write these files* may be better as one module with settings than as thirty-five modules. Left
open deliberately; it is a question about the shape of the catalogue, not about whether to have one.
## Accepted, 2026-09-29, against what was built
*Marked in a grooming pass: the mesh was built to this record and the record still said `proposed`.*
The proposal is the arrangement that exists. `mesh-catalog` holds descriptions of software we did
not write and the programs that provision it, and holds neither the mesh's own components nor an
application's own module. The mesh's list of modules is a table in the control plane, filled by
`module add`, and every module records the source it came from with the commit it was read at.
**One half is not built: `module check` as a command on the control plane's binary.** A manifest is
still validated by a test that reaches into the control plane's internals — which works for this
catalogue and gives nothing at all to somebody describing their own application in their own
repository, which this record says is the case that matters most. That is
[issue 148](../04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md).
@@ -1,6 +1,6 @@
---
topic: what runs on it
status: proposed
status: accepted
date: 2026-09-25
deciders: jochen
reconstructed: false
@@ -1,6 +1,6 @@
---
topic: what runs on it
status: proposed
status: accepted
date: 2026-09-25
deciders: jochen
reconstructed: false
@@ -275,3 +275,11 @@ On acceptance, each of these is superseded or amended by this record, not edited
everything a module needs is a requirement
- [Issue 095](../04-ISSUES/095-a-module-assigned-after-genesis-has-no-broker-account/00-report.md),
[issue 103](../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md): what fails today
## Accepted, 2026-09-29, against what was built
*Marked in a grooming pass.* The vault is a module providing `secret` at mesh scope, and six
modules in the catalogue require it — so a shared secret is a requirement answered by the vault,
which is what this record asks for. Private keys are still made where they are used and never
travel, which is the other half and was never in question.
@@ -1,6 +1,6 @@
---
topic: what runs on it
status: proposed
status: accepted
date: 2026-09-26
deciders: jochen
extends: 0112-a-module-definition-names-no-node-mesh-or-path.md
@@ -47,3 +47,10 @@ other boundary already is: the module name.
node runs one of each (ADR 0115)" — instead of failing on whichever name collides first.
- Multi-tenant asks are answered in the catalogue (a second module definition), not in the
control plane.
## Accepted, 2026-09-29, against what was built
*Marked in a grooming pass.* The rule is enforced where it cannot be forgotten: `assignment`'s
primary key is `(node, module)`, so a second assignment of one module to one machine is not a thing
the mesh can hold. The record read `proposed` while the schema had already settled it.
@@ -12,7 +12,7 @@ extends: 0102-the-mesh-writes-into-a-shared-file-never-over-it.md
## Context
When a resource stops being declared — its module unassigned, the node sent a
deliberately-empty declaration ([issue 127](../04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)),
deliberately-empty declaration ([issue 149](../04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)),
or a new catalogue version renaming its id — the host undoes it. The host's own code states
the rule it means to follow: **it removes what it made and leaves what it merely configured.**
For almost every resource it does exactly that:
@@ -0,0 +1,150 @@
---
topic: the mesh
status: accepted
date: 2026-09-28
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md
---
# 132. A seat carries the tools its holder must serve
## Context
[ADR 0129](0129-a-seat-carries-the-protocol-of-its-role.md) gave a seat the protocol of its role in
three parts: the work it accepts, the events it emits, and the verbs it **serves** — request and
reply, awaited. The bus already derives authority from all three: a holder subscribes
`mesh.seat.<seat>.tool.<verb>`, and a module that uses the seat may publish it and nothing else.
**The serving third has never been used.** The mesh defines 14 seats, 8 mesh-scoped and 6
node-scoped. Exactly one carries a protocol at all — the build machine, which accepts `build` and
emits `built`. Not one seat declares a single verb it serves. The mechanism is built, enforced, and
empty.
Meanwhile every tool on the mesh is addressed to a module. Of 72 modules in the catalogue, 45 serve
tools, about 203 of them, each on `mesh.mod.<module>.tool.<name>`. So a caller binds to the module
that happens to hold a role rather than to the role, and replacing that module breaks every caller —
which is the thing seats exist to prevent everywhere else.
**Nothing can say what tools exist.** Measured on 2026-09-28, with the bus carrying the whole mesh: a
workstation client holding an operator credential connected, the bus accepted the account, and
`mesh call gitea.gitea_list_repos` answered with real repositories. The same client's `mesh tools`
found nothing, because it asks `mesh-catalog.catalog_tools` and no module serves that: the catalogue
serves `catalog_modules`, `catalog_module`, `catalog_provides`, `catalog_dependents` and
`catalog_stale`. An agent can therefore call any tool it already knows the name of and discover none.
MCP's `tools/list` is that same question, so the MCP surface is a working transport over an empty
catalogue.
**And there is nowhere for a tool's definition to live.** A manifest has a `tools` field: 0 of the 45
modules that serve tools fill it. That is not neglect, it is the arrangement failing — the field was
the bus grant's source for what a module may subscribe, and because nothing filled it every module
that served a tool was refused its own subscription on the new bus, live, until the grant was changed
to the module's own namespace. Today a tool's name, description and argument schema exist only in the
module's code.
Two facts about the machinery matter for what follows. A seat's protocol is not in the store: the seat
rows lack the ADR 0129 columns, so the protocol comes from compiled defaults and is merged in when a
row is read. And `seatSubject` is flat — `mesh.seat.<seat>.<kind>.<verb>` with no node in it — so a
node-scoped seat's tool call would reach every node's holder at once, and the holders' queue group
would hand it to whichever answered first.
## Decision
**A seat's protocol carries its tools in full**: the verb, what it does, and the schema of its
arguments and of its answer. The seat is the definition of the role's interface; the holder is an
implementation of it.
**Serving the seat's tools is a condition of holding the seat.** A module that does not serve every
verb the seat declares may not occupy it. This is checked where the other conditions of holding are
checked — registration and handover — and refused by naming the verbs that are missing.
**A role's tools are addressed to the role.** `mesh.seat.<seat>.tool.<verb>` mesh-wide. A node-scoped
seat carries the node in the address, because one subject reaching six machines' holders is not an
address, and the queue group that made it look like one would silently pick a winner.
**A module keeps its own tools, and both exist.** `gitea_list_repos` stays, because gitea can run
without holding the `git` seat — a second forge, an instance kept for one purpose. The module's name
answers *this gitea*; the seat's verb answers *whoever is the forge*. Which of the two a caller wants
is a decision in the running session, not one the mesh makes for it.
**What answers "what tools exist" follows where the definition lives.** A seat's tools are read from
the mesh's own records. A module's own tools are answered by the module, from the code that defines
them. Discovery is therefore a read for the durable half and a question to the running mesh for the
free half.
**A seat's tools are an interface, and change like one.** Additive within a version; a change that
would break a caller takes the version token the subject already has room for (design 29 §8), and the
two run side by side until nothing is bound to the old one.
**The mesh's own verbs are the `mesh-controller` seat's tools.** `status`, `push`, `build`, `assign`
and the rest are a role's interface, not a container's, and the audit point [ADR 0095](0095-the-control-plane-is-the-way-to-ask-a-module.md)
asks for is the seat's holder.
## Options considered
1. **The manifest declares each module's tools.** Rejected. The list is then written twice — in the
manifest and in the code — and a schema in a manifest goes stale silently, which is the worst kind
of wrong for something an agent reads to decide what to call. It is also the arrangement that has
already failed once: the field exists, 0 of 45 modules fill it, and the grant that depended on it
refused every tool subscription on the mesh.
2. **Every runtime answers an introspection call, and something aggregates them.** Rejected as the
shape for a role's tools, kept for a module's own. An aggregator needs permission to publish into
every module's namespace, which is a widening the mesh otherwise gives only to the control plane;
and the answer is only as available as the modules are, so a mesh whose catalogue cannot say what a
role answers while its holder is down cannot plan against it.
3. **The control plane answers everything.** Rejected. It puts a tool surface on the control plane for
tools it does not implement, and makes discovery depend on the one component that must stay
answerable while it is itself being replaced. The mesh's own verbs are its to answer, and it answers
them as the holder of a seat.
4. **Seats only; no module tools.** Rejected. Most modules hold no seat, and inventing a seat per
module to give its tools a home would dilute what a seat is: one holder of a role the mesh needs
exactly one of.
## Consequences
**One capability can have two names, deliberately.** A forge that holds the `git` seat answers both
`mesh.seat.git.tool.list_repos` and `mesh.mod.gitea.tool.gitea_list_repos`. This is the one place the
mesh accepts two names for one thing, because they are answers to different questions and the second
one survives the module not holding the seat. The glossary rule stands everywhere else.
**A seat becomes a contract to implement.** Adding a verb to a seat is a change every holder must
make, and a claim that was valid becomes invalid until it does. That is the point, and it is also the
reason a seat's tools should be few and durable while a module's own stay free.
**Three prerequisites, none of them in place.** The seat's protocol must be in the store rather than in
compiled defaults, or discovery reads a binary rather than the mesh. The protocol must become richer
than a list of verbs, because a verb without a schema is not something an agent can call. And a
node-scoped seat needs the node in its subject before any of its tools can exist.
**Discovery becomes cheap for the half that matters.** What roles the mesh has and what each answers is
a query, with no fan-out and nothing to be up. An agent's authority can then be role-shaped — *the
forge's tools* — rather than a list of module-specific names that changes when a module is replaced.
**The MCP surface belongs inside the mesh.** Once the tools are the mesh's own records, the thing that
serves them to an agent is a module the mesh assigns to the machine where the agent sits, with a
credential the mesh minted and authority derived from what it may call — not a program started by hand
with a credential printed to a terminal.
## How this is checked
- **Holding is refused without the verbs.** The condition sits with the other conditions of holding a
seat, so registration and a handover both refuse a module that does not serve what the seat declares,
and the refusal names the missing verbs. A test per condition, as the other seat conditions have.
- **The grant is derived from the seat, and already is.** A holder's subscription and a user's publish
come from the seat's protocol, so a verb nobody declared is a subject nobody may use, and a verb the
seat declares reaches exactly its holder. The golden composition of the bus's user list is the test
that keeps it honest.
- **Discovery is a read, and is tested as one.** What the mesh answers for a seat's tools equals what
the seat's records declare — no call to a module in the path, so the test needs no running module.
- **A node-scoped seat's subject carries its node**, checked by the same test that checks the subject
table: two nodes holding one node-scoped seat derive two addresses.
## References
- [ADR 0129](0129-a-seat-carries-the-protocol-of-its-role.md) — the protocol this widens
- [ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md) — the role, its work and its events
- [ADR 0095](0095-the-control-plane-is-the-way-to-ask-a-module.md) — a tool call passes one process where an audit belongs
- [ADR 0126](0126-a-module-declares-its-own-seats.md) — an event is addressed to its emitter, for the same reason a role's verb is addressed to its role
- [`03-DESIGN/01-to-be/26-the-seats.md`](../03-DESIGN/01-to-be/26-the-seats.md) — how a seat is held and handed over
- mesh-controller #116, #117, #118 — the grants as they now stand: a module serves its own namespace, the control plane may ask any tool
- Measured 2026-09-28 on the live mesh: an operator credential calling a module's tool over the bus answers; `tools/list` finds nothing
@@ -0,0 +1,159 @@
---
topic: what runs on it
status: superseded
date: 2026-09-28
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md
superseded-by: 02-DECISIONS/0135-a-module-version-prepares-its-state-before-it-runs.md
---
# 133. A module owns its migrations, and the mesh owns when they run
## Context
On 2026-09-28 the mesh replaced its own control plane, through its own upgrade path, with a build
carrying a migration. Nothing applied it. For the next three quarters of an hour every build the mesh
made was refused by the store with one line — *column "built_contexts" does not exist* — which reached
only whoever happened to be waiting on that build's reply. The images were built and published, so the
registry filled with artifacts the mesh has no record of, and the overview went on reporting that every
module was current ([issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md)).
The schema had been created once, at genesis, by an action in the foundation bundle. Nothing ran it
again, through many updates of the control plane since.
**The mechanism to do this right already existed and one module used it wrong.**
[ADR 0052](0052-a-step-that-runs-once-before-a-container.md) makes a run-once container a step the host
runs to completion before whatever the declaration places after it, and names migrating a schema as the
case it exists for. Three facts about how it is used today:
- The control plane's manifest had no step at all. The immediate fix was to write one by hand, and that
hand-written step repeats three environment variables and three volume mounts from the server
resource it precedes — six chances to drift from the thing it prepares.
- Two other modules hand-write the same shape for the same reason: gitea's admin bootstrap and
mosquitto's dynsec seed, each repeating its sibling's image, environment and mounts. One of them
ends in `|| true`, which is a lock implemented as a shrug.
- The catalogue module takes the other road: it migrates its own schema in its own code when it starts.
That failure mode is a crash loop rather than a stop — the catalogue restarted 338 times this
morning on an unrelated start-time failure, and nothing anywhere said the mesh's graph had a gap.
**What the mesh already has, and what HAL needed stages for.** Ordering a provider before its consumer
is `providersFirst`, which topologically orders a node's modules. Ordering within a module is
declaration order, and a run-once container gates everything after it. Remembering that a step has
already run is the digest of its declaration, recorded only after it exits 0
([ADR 0018](0018-a-picture-is-read-from-what-runs.md)) — and the image is part of that digest, so a new
build re-runs it. Three of the four things a stage system provides are therefore already here. The
fourth — that a module has a schema at all — is the only thing missing.
**Nothing in the catalogue ships a migrations directory.** Of 72 modules, none has one; the modules that
migrate do it in their own code. So this is not a decision about where SQL files live. It is a decision
about who runs them and when.
**Two facts bound what is safely expressible.** A node converges toward its own declaration without
waiting on any other node. And of the five modules that run on more than one machine today — dnsmasq,
fail2ban, networking, networkmanager, sshd — not one wants a store; every module with a database is on
exactly one machine.
## Decision
**A container may declare steps to run before it.** The same container, run to completion, with
different arguments, in order, before it starts. The mesh derives the run-once resources from that
declaration, so the image, the environment, the volumes, the network and the credentials come from the
one place they are already described and cannot drift from it.
**A module's migrations are the first user of this, and the module owns them entirely.** The SQL, the
order, the idempotence, the lock, and which dialect it speaks. The mesh never learns that postgres and
mssql differ, because it runs the module's own image with the module's own arguments against the
module's own binding and requires exit 0. A module needing both stores runs one step that does both.
**The mesh owns the moment, and the gate is the guarantee.** Whether a version may serve when its
schema is not there yet is a deployment question, and the mesh is the only thing that can answer it,
because the mesh is what starts the container. A step that fails stops the container it precedes, so
a failed migration is a version that does not serve rather than a version serving against a store it
does not match.
**Per node, and there is no level.** The step runs wherever the module runs. A step that ran "once,
somewhere" would leave every other machine with no gate at all, and additive migrations protect old
code against a new schema, never new code against an old one. The cost is an obligation a migration
runner already carries: a version table and a lock.
**"Once, mesh-wide" is what holding a seat means.** A step that is not idempotent — seeding an
account, sending a notice, taking a backup — belongs to a module that holds a seat, where the mesh
already guarantees one holder, on record, handed over deliberately. That is the answer to the level
question rather than a field that has to invent an election and keep it somewhere.
**Migrations are forward-only and additive.** The step runs before the *new* container starts, so the
old one is still running against the new schema for the length of the apply.
**Declared, never inferred.** The control plane cannot see inside an image, so a module that ships
migrations and declares no step is not refusable at registration; it breaks on its first upgrade. This
record says so rather than implying a check that cannot exist.
## Options considered
1. **Each module migrates itself when it starts** — what the catalogue does today. Rejected: it turns a
schema failure into a crash loop instead of a stop, it is invisible in the declaration so nothing can
say the module even has a schema, and two machines running the module both migrate at start with
nothing sequencing them.
2. **The mesh applies migrations itself**, with a driver and a version table per store — HAL's shape.
Rejected: the mesh would have to know one store type from another, hold another module's store
credentials, and reach a machine to use them, which [ADR 0005](0005-the-node-host.md) forbids. It is
also the reason that shape needs levels: something central has to decide where the once happens.
3. **A hook lifecycle** — pre-build, post-build, pre-deploy, post-deploy. Rejected: there is no deploy
event here to hook. A declaration is a desired state applied in order and reconciled forever, so
"pre-deploy" is exactly "a step before this container", pre- and post-build are what a Dockerfile and
the artifact list already are, and "post-deploy" has no moment to name.
4. **A hook level** — once per module, or once per module-node assignment. Rejected as a field, kept as
a property: see the decision. A once-per-module step needs cross-node ordering underneath it to be
safe, and a node converging without waiting on its neighbours is worth losing on purpose rather than
by accident.
5. **Every module hand-writes its own run-once step** — the immediate fix for the control plane.
Rejected as the general answer: it duplicates the resource it precedes, in three places already, and
a hand-written step is one the next module forgets. Forgetting it is the fault this record exists
for.
6. **Record a schema level per module in the store.** Rejected: gating makes the invariant true by
construction, so a level is a second account of the same fact and the first one to go stale.
## Consequences
**Three hand-written steps collapse into one line each**, and the control plane's own migrate step stops
repeating its server's environment and mounts.
**The catalogue's self-migration becomes the exception to remove.** One shape, and the mesh's own
control plane is not an exception to it either.
**A module on two machines with one shared store must lock.** Today none is, so this is an obligation
stated before it is needed rather than discovered by two concurrent migrations.
**There is still no readiness-gated step.** Only an action carries `verify`; a container has no health
notion, so "run this once the service answers" remains unexpressible and seeding through a running
service's API has no home. That is its own decision about a container's readiness, and this record does
not make it.
**Genesis keeps its own action.** At birth there is no control plane to derive anything from, which is
what [ADR 0067](0067-genesis-is-a-pivot.md) already says about that moment.
## How this is checked
- **The composition carries the step.** A test on a node's composed declaration: every container that
declares steps before it is preceded by them, and the derived step's image, environment, volumes and
network equal the container's — so the two cannot drift, which is the failure the hand-written kind
has.
- **A failed step stops what follows.** The host already refuses to go on past a run-once step that did
not exit 0; the test for that is extended to a derived one, so the gate is checked rather than
assumed.
- **The mesh's own schema is covered by the same mechanism as everything else.** The control plane
declares its step in its own manifest, so the case that failed on 2026-09-28 is the case the test
covers.
- **A module claiming a seat for a once-only step is checked where seats are checked** — the conditions
of holding, not a new mechanism.
## References
- [ADR 0052](0052-a-step-that-runs-once-before-a-container.md) — the step this extends
- [ADR 0018](0018-a-picture-is-read-from-what-runs.md) — a digest is the record that something happened
- [ADR 0005](0005-the-node-host.md) — the control plane decides and never touches a machine
- [ADR 0067](0067-genesis-is-a-pivot.md) — why genesis does it differently, once
- [issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md) — the failure that produced this record
- [`03-DESIGN/01-to-be/32-what-a-module-declares.md`](../03-DESIGN/01-to-be/32-what-a-module-declares.md) §6 — the lifecycle this sits in
- Measured 2026-09-28: three hand-written run-once steps repeating their sibling's resource; 0 of 72 modules with a migrations directory; 5 modules on more than one machine, none of them wanting a store
@@ -0,0 +1,126 @@
---
topic: the mesh
status: accepted
date: 2026-09-28
deciders: jochen
reconstructed: false
---
# 134. The mesh says what it applied
## Context
The pipeline is observable on the bus from a merge to an artifact, and modules already plug into it:
the forge emits `pull.merged`, the build machine's seat emits `built`, the catalogue emits `registered`,
`upgraded` and `rebuild-needed`, providers emit `postgres.database.provisioned` and its siblings. Things
consume them today — the catalogue consumes `built`, each provider consumes its own provisioning events,
model-usage consumes `*.usage.*`, the audit logger consumes `**`. Nothing had to be invented for any of
that; subscribing *is* plugging in.
**It goes dark at the moment it touches a machine.** A host applies a declaration and reports to the
control plane on the control branch, which only the control plane may read — correctly, because a report
carries what a machine is and enrolment travels the same way. So nothing on the mesh says *this machine
now runs version Y of module Z*, or that it refused to, or why.
What that cost on 2026-09-28, in one morning:
- A build result the store refused was visible only to whoever was waiting on that build's reply. For
three quarters of an hour the mesh built things and recorded none of them, while the overview said
every module was current ([issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md)).
- A module crash-looping at start — 338 restarts — was found by reading a container's logs by hand.
Nothing on the bus said the mesh's graph had stopped learning.
- A run-once step that fails now stops an upgrade by design ([ADR 0133](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md)),
and the same silence would cover it: the version simply would not appear.
**And the one thing the control plane does emit is refused by its own account.** Answering a catalogue
that asks to catch up, it publishes each recorded build under `mesh.mod.control-plane.event.…` — a
module namespace for a module that does not exist. Its own permissions refuse it, so a catalogue that
restarts gets nothing and keeps its gap. The control plane has facts to state and nowhere to state
them.
## Decision
**The mesh emits the deploy half of the pipeline as facts on the bus.** What a machine now runs, and
what it refused to run and why. Both are facts about the mesh doing its work, in the same form as every
other fact on the bus, so anything that wants them subscribes the way the catalogue subscribes to
`built`.
**The control plane states them, as the holder of the `mesh-controller` seat.** Its facts live under the
seat's own namespace, which is where a role's events belong
([ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md),
[ADR 0129](0129-a-seat-carries-the-protocol-of-its-role.md)) and which survives the control plane being
replaced. That is also what gives the catch-up replay a subject it may publish instead of an invented
module namespace.
**Emitted when what a machine runs changes, not on every convergence pass.** A host reconciles
continuously and reports each time; a fact per pass would be a fact per minute per machine that says
nothing. The report carries the declaration it applied and what changed, so the control plane has what
it needs to speak only when there is something to say.
**A refusal is a fact with a subject in it** — which machine, which resource, and the reason as the host
gave it. A refusal that names only the machine is the silence this record is about, one level up.
**Reports stay where they are.** A node's report remains control traffic that only the control plane
reads. The deploy facts are derived from it, which makes them second-hand on purpose: one emitter, one
ordering, and no widening of the narrowest account in the mesh.
## Options considered
1. **Leave it as it is, and let whatever cares ask the control plane.** Rejected: asking for a fact that
already arrives is the shape the mesh removed everywhere else, and nothing can react at the moment a
machine changes — which is exactly when a graph, an audit or an operator wants to know.
2. **Each node emits its own facts.** Rejected: it widens every host's account to an event namespace, and
a host's authority is deliberately the narrowest in the mesh. Its report already reaches the one thing
that can speak for it.
3. **Widen who may read the control branch.** Rejected: that branch carries what machines say *to* the
control plane, enrolment included. Widening its readers widens that too, for an unrelated reason.
4. **A registry of deploy hooks** — something registers interest and is called. Rejected: an event is
already the mechanism; there is nothing to register, and a callback is an address the mesh spent
[issue 102](../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)
learning not to keep.
5. **Put the facts on the node's own declaration stream.** Rejected: that stream is last-per-subject by
design, so a machine away for an hour gets exactly the current declaration and nothing older. A
history of what happened cannot live in a stream built to forget.
## Consequences
**The audit logger gets the deploy half for nothing**, because it consumes everything.
**A failure becomes visible where the mesh is watched** rather than where someone happened to be
looking. That answers the open question [issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md)
left about a record the store refused.
**The catch-up replay stops being a burst of events.** With the control plane able to state its own
facts, replaying history as if it were happening now is a choice rather than the only option — and the
better shape is the question the catalogue is actually asking, answered once
([design 33](../03-DESIGN/01-to-be/33-the-tools-the-mesh-answers.md)).
**The facts are second-hand.** The control plane says what a machine reported, so a machine that cannot
reach the bus produces no fact at all. Absence is not health, and what a machine was last heard from
stays the place that says so.
**The events stream carries more.** Bounded by emitting on change rather than on every pass, and each
fact is small; the stream's own limits remain what keeps it finite.
## How this is checked
- **The control plane's grant names exactly the subjects it emits**, derived from its seat like every
other principal's, and the composed user list is compared against a golden file — so a fact it cannot
publish fails a test rather than a catalogue's replay.
- **A convergence that changed nothing emits nothing.** A test with two identical reports and one
expected fact, because the failure this guards against is a fact per minute per machine.
- **A refusal names its resource.** A test where a host reports a failed resource and the emitted fact
carries which one and why, not merely that something went wrong.
- **What the mesh emits is what something consumes.** The subject a module declares it consumes derives
to the subject the control plane publishes — the same agreement test that already keeps the
controller's own subscriptions honest.
## References
- [ADR 0041](0041-events-are-a-relationship.md) — an event is a relationship, not a call
- [ADR 0126](0126-a-module-declares-its-own-seats.md) — an event is addressed to its emitter, because the emitter's identity is the meaning
- [ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md), [ADR 0129](0129-a-seat-carries-the-protocol-of-its-role.md) — a role's events belong to the role
- [ADR 0083](0083-one-push-leaves-the-mesh-consistent.md) — a report is held for the store rather than lost
- [ADR 0133](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md) — the gate whose failure this makes visible
- [`03-DESIGN/01-to-be/32-what-a-module-declares.md`](../03-DESIGN/01-to-be/32-what-a-module-declares.md) §6 — the lifecycle, which ends today at a report nobody else may read
- Measured 2026-09-28: 45 minutes of builds recorded nowhere with the overview reporting health; a module at 338 restarts found by hand; the control plane's only emitted event refused by its own permissions
@@ -0,0 +1,153 @@
---
topic: what runs on it
status: accepted
date: 2026-09-28
deciders: jochen
reconstructed: false
supersedes: 0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md
---
# 135. A module version prepares its state before it runs
## Context
[ADR 0133](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md) settled who runs a
module's migrations and when, and it said so in the wrong vocabulary. It put the declaration on a
*container* — "a container may declare steps to run before it" — and derived the scope of the work from
the *machine*. Both are wrong at the level a module author works at, and the second is wrong on the
facts.
**A container is one resource kind the host applies.** A module has code, state and a version; whether
its artifact is an image, a bundle or something later is the mesh's business. The module-facing
vocabulary for a module's own code already exists and has nothing to do with a container runtime: a
module declares **entrypoints** — this file is my tools, this file is my provisioner — and the mesh runs
them. A manifest that says "run this container with these arguments, and here are the volumes and
environment again" has an author writing down the machine's business twice.
**And the scope is not the machine's to decide, because the mesh already decided what a state is.** A
consumer is a module *on a machine* (migration 0015, from
[issue 022](../04-ISSUES/022-one-credential-per-node-per-provision-not-per-module/00-report.md)):
the mesh derives a login per consumer and the provider creates a database owned by exactly that login
([ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)). So a module on three machines is three
consumers, three credentials and three databases. There is no shared state for two machines to race over,
and ADR 0133's central caveat — that a module's migrations must take a lock because two machines might
migrate at once — describes a situation the mesh does not currently produce.
That correction makes the whole "level" question HAL answered with stages disappear: the scope of
preparation is the scope of the state, and the mesh knows it.
What the earlier record got right and this one keeps: the module owns the work, the mesh owns the moment,
the gate is the guarantee, migrations stay forward-only, and none of it can be inferred from inside an
artifact. What produced it also stands — the control plane was replaced with a build carrying a migration,
nothing applied it, and for three quarters of an hour every build was refused by the store with one line
that reached only whoever was waiting on a reply
([issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md)).
## Decision
**A module version declares an entrypoint that prepares its state.** One name in the manifest, in the
same vocabulary as the entrypoints it already declares for its tools and its provisioner. No container,
no command line, no environment, no mounts — those are how a machine runs the module's code, and the
module already said that once.
**The mesh runs it as it runs that module's own code, to completion, in the module's own context.** Every
binding, credential and setting the module's code would receive, because it *is* the module's code. How a
machine does that is the host's business and stays there: for an image artifact it is the step
[ADR 0052](0052-a-step-that-runs-once-before-a-container.md) already defines, and a later kind of artifact
changes the host, not the manifest.
**Preparation gates the version.** A version whose preparation did not succeed does not run — anywhere.
Since the rollout already sends machines one at a time and stops at the first that does not take a
version, a preparation that fails stops the rollout there, leaving every other machine on the version
that works.
**Preparation is scoped to the state, and the mesh derives that scope.** State the mesh provisions is per
consumer — a module on a machine — so preparation happens once per consumer. State the module keeps on
the machine is per machine, which is the same answer. A module that holds an exclusive seat has one of
itself, so its preparation happens once by definition. No level, no election, no cross-node ordering, and
no lock obligation invented for a race the mesh does not create.
**Once per version per state.** A version bump attempts preparation once against each state it has; the
module's own runner decides there is nothing to do, which is what a runner with a version table does
anyway. A retry after a partial failure runs it again, so the work is the module's to make safe against
that — the one obligation no design can remove.
**Forward-only and additive.** Preparation runs while the previous version is still serving, so a
migration that removes or renames what the old code reads breaks the mesh in the window between the two.
**Declared, never inferred.** The control plane cannot see inside an artifact, so a module that ships
migrations and declares no entrypoint is not refusable at registration. It breaks on its first upgrade,
and this record says so rather than implying a check that cannot exist.
## Options considered
1. **A container declares steps before it** — [ADR 0133](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md).
Superseded, not because the mechanism is wrong but because the *declaration* is in the wrong place: it
makes every module author restate the machine's arrangement, and it ties a module's own lifecycle to
one resource kind. The host-side mechanism it named is retained and is now an implementation detail.
2. **Each module prepares itself when it starts** — what the catalogue does today. Rejected: a schema
failure becomes a crash loop rather than a stop, nothing in the declaration says the module has a
state to prepare, and the version serves the moment it starts rather than after the state is right.
3. **The mesh applies migrations itself**, with a driver and a version table per store type. Rejected:
the mesh would have to know one store from another, hold another module's credentials and reach a
machine with them, which [ADR 0005](0005-the-node-host.md) forbids. It is also what forces a stage
system: something central has to decide where the work happens.
4. **A hook lifecycle** — pre-build, post-build, pre-deploy, post-deploy. Rejected: a declaration is a
desired state reconciled forever, so there is no deploy moment to hook. "Pre-deploy" is exactly this
record; pre- and post-build are what a recipe and the artifact list already are; "post-deploy" names
nothing that happens.
5. **A declared level** — once per module, or once per assignment. Rejected: the mesh already knows what a
state is, so asking an author to choose is asking them to restate a fact the mesh holds, with a chance
of contradicting it.
6. **Record a preparation level per module in the store.** Rejected for the reason ADR 0133 gave and this
record keeps: gating makes the invariant true by construction, and a level is a second account of the
same fact.
## Consequences
**An author's whole contract is one line, once.** Write the migration in the module's code, name the
entrypoint that runs it, and every later version rolls out as: build, prepare, run — with nothing
per-version to remember and nothing about the machine to restate. That is the property this exists for.
**Three hand-written steps in the catalogue collapse**, and the control plane's own migrate step stops
repeating its server's environment and mounts.
**The catalogue's self-preparation becomes the exception to remove.** One shape, and the mesh's own
control plane is not an exception either.
**A module scaled across machines with one shared state is not expressible**, and this record does not
make it so. The mesh gives each consumer its own state; a deliberately shared one is a different
provision model, and the place the "once, mesh-wide" question would genuinely return. Named here so it is
a decision when it happens rather than a surprise.
**There is still no readiness-gated step.** Only an action carries `verify`; nothing declares that a
service answers, so preparation that must happen *after* something is serving — seeding through its own
API — remains unexpressible.
**Genesis keeps its own action.** At birth there is no control plane to derive anything, which is what
[ADR 0067](0067-genesis-is-a-pivot.md) says about that moment.
## How this is checked
- **The composition carries the preparation, in the module's own context.** A test on a node's composed
declaration: a version declaring a preparation entrypoint is preceded by it, and what it is given
equals what the module's own code is given — asserted equal rather than written twice, which is the
drift the superseded shape invited.
- **A preparation that fails stops the version.** The host does not go past a step that did not complete,
and the rollout stops at the first machine that did not take a version. Both are existing behaviours
with existing tests; the test for preparation asserts the two together — the machine does not run it,
and the machines after it are left alone.
- **Once per version per state.** A test that a second convergence of the same version prepares nothing,
and that a new version prepares again.
- **The mesh's own control plane declares one.** The case that failed on 2026-09-28 is the case the tests
cover, rather than a case a comment says is covered.
## References
- [ADR 0133](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md) — what this supersedes, and why
- [ADR 0052](0052-a-step-that-runs-once-before-a-container.md) — the host-side step that implements it for an image artifact
- [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md), [issue 022](../04-ISSUES/022-one-credential-per-node-per-provision-not-per-module/00-report.md) — a consumer is a module on a machine, which is what makes the scope derivable
- [ADR 0018](0018-a-picture-is-read-from-what-runs.md) — a digest is the record that something happened
- [ADR 0005](0005-the-node-host.md) — the control plane decides and never touches a machine
- [ADR 0134](0134-the-mesh-says-what-it-applied.md) — what makes a failed preparation visible
- [issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md) — the failure that produced both records
@@ -0,0 +1,106 @@
---
topic: what runs on it
status: accepted
date: 2026-09-28
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md
---
# 136. A step gates its module, not the machine
## Context
[ADR 0052](0052-a-step-that-runs-once-before-a-container.md) made a run-once container a step the host
runs to completion, and gave it the same reach a failed action has: it stops everything the declaration
places after it. When the only steps on the mesh were a broker's seed and a forge's admin account, that
reach was invisible — the thing after the step was the container the step existed for, in the same
module.
[ADR 0135](0135-a-module-version-prepares-its-state-before-it-runs.md) made a step something the mesh
derives for **any** module that prepares its state, and that turns the reach into a fault. A module
whose database is briefly unreachable now stops every module declared after it on that machine, for as
long as it is unreachable.
**The host already rejected this for every other shape, and says why in its own loop.** From
[issue 011](../04-ISSUES/011-one-broken-module-blocks-every-other/00-report.md):
> It used to stop at the first one, and that made one broken resource hold the whole machine hostage: a
> module declaring a package that does not exist meant every module ordered after it was never applied,
> for ever, and the mesh reported "failed" without saying that the rest had not been tried. A machine
> with one bad module and nine good ones ran none of the nine.
Everything is attempted and every failure reported — except an action and a run-once step, kept as the
deliberate exceptions. So the mesh has two rules about the same question and the wider one is now
reachable by any module that declares a schema.
**And it deadlocks a case the catalogue already named.** The catalogue migrates its own schema when it
starts rather than in a step, and says why in its code: *a schema step that had to reach the provider
over the overlay would block the very apply that brings the overlay up*. With a machine-wide gate that
is exactly right — the step fails, the apply stops, the overlay module after it is never applied, and
the next reconcile is blocked the same way. The module that most obviously wants a step could not have
one.
## Decision
**A step gates its own module.** A run-once container that does not complete stops the rest of *that
module's* resources and nothing else. Every other module on the machine is attempted, as every other
shape already is.
**An action still gates the machine.** Genesis is a row of actions, each making the next possible, and
they belong to no module — there is nothing narrower for their reach to be.
**What was not attempted is reported, not inferred from silence.** A skipped resource appears in the
machine's account of the apply as skipped, with the reason, because "not attempted" and "nothing to do"
are different answers and only one of them is somebody's to fix.
**A module is the part of a resource's identity before the first dot**, which is how the mesh composes
them. What the mesh declares in its own right — a guard, an opening, the adoption's own resources —
belongs to no module, and its gate is therefore the machine's.
## Options considered
1. **Leave the reach as it is.** Rejected: it reintroduces, through a mechanism now derived for every
module, exactly the fault issue 011 removed. A mesh where one module's unreachable database stops a
machine converging is worse than one where that module alone is behind.
2. **Make preparation not a gate at all** — run it and carry on. Rejected: then a version serves against
a state nobody shaped, which is the whole of what ADR 0135 exists to prevent.
3. **Order every module's step before everything else on the machine**, so a gate stops nothing that
matters. Rejected: it inverts the order a module needs — its files and directories are declared before
its step because the step reads them — and it would still stop later modules.
4. **Let a module declare how far its step reaches.** Rejected: the answer is the same for every module,
and a field would let one be wrong about it.
## Consequences
**The catalogue can move to a step.** The reason it migrates at start — that a step blocks the apply
that would make its provider reachable — stops being true: the step fails, that module waits, the
overlay comes up, and the next reconcile prepares it. One shape for the whole mesh, which is what
ADR 0135 asked for and could not have had.
**A module can sit behind while the machine is otherwise current.** That is the honest state and it is
what the report now says. It also means a preparation that never succeeds is a module that never
upgrades, quietly, until somebody reads the report — which is an argument for
[ADR 0134](0134-the-mesh-says-what-it-applied.md) rather than against this.
**A module's resources must be ordered within the module for the gate to mean anything.** They already
are: the mesh composes a module's resources in the order its manifest declares them, and its own
workload comes after the files it reads.
## How this is checked
- **A failed step stops its module and nothing else.** A test with two modules: the one whose step
failed does not start its workload, the other starts, and the error still says the failure gated
something. It fails against the previous behaviour, which is how it was written.
- **An action still stops the machine.** The existing test for a failed action is unchanged, and a step
with no module in its identity — which is what genesis carries — takes the same path.
- **The report names what was skipped.** Asserted in the same test, because a gate nobody can see is
indistinguishable from a module that had nothing to do.
## References
- [ADR 0052](0052-a-step-that-runs-once-before-a-container.md) — the step this narrows
- [ADR 0135](0135-a-module-version-prepares-its-state-before-it-runs.md) — what made the reach reachable
- [issue 011](../04-ISSUES/011-one-broken-module-blocks-every-other/00-report.md) — the same fault, removed once already
- [ADR 0134](0134-the-mesh-says-what-it-applied.md) — how a module left behind becomes visible
- mesh-host `internal/apply` — the loop whose own comment argued this case for every other shape
@@ -0,0 +1,117 @@
---
topic: what runs on it
status: superseded
date: 2026-09-28
deciders: jochen
extends: 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
reconstructed: false
superseded-by: 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md
---
# 137. A machine says which networks it routes
## Context
The filter the mesh derives denies forwarding by default, because without a forward chain it says
nothing about a container's published port
([ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md), [issue 047](../04-ISSUES/047-the-firewall-does-not-cover-published-container-ports/00-report.md)).
To keep a machine's own containers working it then allows two ranges: the container runtime's
default bridge pool, and the pool its compose files are given. Those two are named in the
controller's code, with a comment saying what the gap is:
> A machine whose runtime is configured with something else needs this to say so — which is a thing
> the mesh cannot derive and a reason this list is named here rather than computed.
**There was no way to say so.** The list was a constant. A machine whose guests live anywhere else
was filtered by a rule that looked deliberate and was a guess.
**Measured, on the day a workstation was converged.** Flipping it cut egress for five of its
container networks at once, and for every network its test beds create — the beds allocate a fresh
range per run, from a pool neither default covers. Nothing reported a fault. The containers could
not reach anything, the machine went on reporting that it had applied what it was told, and the
converge preview had said nothing about it either, because the preview lists what *listens* and
routing is not a listener.
**And two questions, not one.** A guest also asks its host for an address and for names. Both arrive
at the input chain, where nothing declared them, so denying by default left the guests of a routed
network with no address and no resolution — which is not a closed port but a network that does not
function, asked for by this machine's own guest.
**Why the machine cannot simply be read.** A test bed creates its bridge while it runs, between one
declaration and the next, so a filter derived from what the machine last reported would be correct
only for the networks that already existed when it was composed. A declared range covers the ones
that do not exist yet.
## Decision
**A machine says which networks it routes for what it hosts, and the filter forwards them.** A
node-level fact, beside the node's public domain
([ADR 0066](0066-public-routing-is-name-agnostic.md)) and for the same reason: the
machine routes them, and the module that loads the filter holds a seat and may be replaced.
**Added to the runtime's defaults, never replacing them.** A machine that names one range has not
stopped hosting whatever was already on the runtime's own pools, and replacing would trade one
silent breakage for another.
**Their guests keep address and name service.** For a network that was named, the input chain admits
that network's own DHCP and DNS, and nothing else: everything else a guest might want from its host
is a port somebody declares, like every other port on this machine.
**Said in CIDR form and checked when it is said.** An entry that does not parse is a line nftables
refuses, and a refused ruleset is a machine filtering nothing while its unit reports a fault — so
the refusal happens where a person can read it, not on the machine.
**A machine that says nothing is filtered exactly as before.** Every machine already converged is
untouched by this.
## Options considered
1. **Leave it constant and edit the code per installation.** Rejected: the value is a property of
one machine, the code is the whole mesh's, and the two ranges as they stand describe a machine
whose runtime was left at its defaults. It is also how this got here.
2. **Derive it from what the machine reports.** Rejected as insufficient, not as wrong: it cannot
cover a network created between two declarations, which is precisely the case that was broken. It
would also make the filter follow whatever appeared on the machine, which is a firewall that
widens itself.
3. **A per-node setting on the module that loads the filter.** Rejected: the machine routes the
networks. The filter module holds a node-scoped seat and is meant to be replaceable, and a
replacement must not lose the machine's own truth.
4. **Replace the defaults with what is said.** Rejected: see the decision. The first machine to name
its bed range would lose its containers.
5. **Admit all input from a routed network, not only address and name service.** Rejected: that is
every port on the machine open to anything it hosts, which is the derivation abandoned.
## Consequences
**The converge preview says what a machine routes**, including when it routes nothing but the
defaults, with the command that changes it. The preview's own sentence about traffic it cannot
preview stays, because a tunnel and the found firewall's NAT are still not previewable.
**A machine whose guests are already broken by an earlier flip is fixed by saying its networks and
pushing**, with no flip to undo.
**The list is one more thing that can be wrong and stale.** A range removed from the machine and
left here keeps forwarding for a network that no longer exists, which admits nothing, because there
is no guest on it to admit. That is the safe direction of being out of date.
## How this is checked
- **What a machine says it routes is forwarded, and its guests keep address and name service.** A
test renders a ruleset for a machine that names one range and asserts both chains, per chain body
so a line in the wrong chain cannot pass it. It fails against the previous behaviour, which is how
it was written.
- **The runtime's own defaults survive naming a range.** Asserted in the same test.
- **A machine that names nothing renders byte-identically to one that names nil**, so every machine
already behind this filter is untouched.
- **Each family is matched in its own syntax.** A test with one v4 and one v6 network asserts
`ip saddr` and `ip6 saddr`, because one set holding both is a syntax error and a ruleset that does
not load is a machine filtering nothing.
- **An entry that is not a network is refused where it is said**, by the parse in the setter.
## References
- [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md) — the derived filter this completes
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — the precedent for a node-level fact
- [issue 047](../04-ISSUES/047-the-firewall-does-not-cover-published-container-ports/00-report.md) — why there is a forward chain at all
- [issue 137](../04-ISSUES/137-converging-a-machine-cut-off-its-own-guests/00-report.md) — the measurement that produced this
- mesh-controller `internal/catalogue/filtering.go` — the constant whose own comment named this gap
@@ -0,0 +1,179 @@
---
topic: what runs on it
status: accepted
date: 2026-09-28
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md
---
# 138. An assignment binds an endpoint and says how far it reaches
## Context
[ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) settled that a machine's
packet filter is derived from what its modules declare they listen on, and that the `from` of a
listen "is the whole of public-versus-internal". That was true of the packet filter, and it turned
out to be true of nothing else.
Reachability is now settled three times, in three places, by three mechanisms that cannot disagree
out loud ([issue 140](../04-ISSUES/140-an-endpoints-reach-is-not-declared/00-report.md)):
- **The filter** reads a listen's source, and a per-node setting may override it. That setting has
exactly one caller in the control plane — the function that builds the node's rules.
- **The names** come from a route contribution, which names a label and a port and says nothing
about reach. The reverse proxy composes a **public** name and an **internal** name for every route
it is given, because it can.
- **The certificate authority** follows from which names exist. Measured on the control-node: an
identity provider carries a public certificate valid 90 days and an internal one valid 24 hours and
renewed daily. No assignment asked for either.
So *this endpoint must not be public* cannot be written. It is therefore enforced by nothing, while a
public certificate for that very name is obtained automatically — the fault
[how-we-build.md](../00-META/how-we-build.md) names, an unenforced rule being indistinguishable from
a wrong one, with the additional cost that the wrong thing is done eagerly.
And a port that is not routed cannot be spoken about at all beyond the filter. The forge serves git
over ssh; that endpoint has no name, no certificate and no way to be called public except a key only
the filter reads.
**Two per-node settings already exist and are half of this.** One gives a module's declared port a
machine port. One overrides a declared port's source. They key on port numbers, so nothing ties a
port to the route that serves it: a route contribution names a port too, and the two are equal only
by coincidence.
**Where this belongs is already decided.** [ADR 0046](0046-a-module-configuration-is-its-assignments-not-its-manifest.md)
says a module's configuration is its assignments. Whether the forge answers git-over-ssh from the
public internet is a fact about one installation and one machine, not a property of the software —
and [ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md) already refuses an
installation's decisions in a definition.
## Considered Options
1. **Leave reach in the manifest, as `from` today.** Rejected: it is an installation's decision
written into the definition, and it cannot differ between two machines running the same module —
which is exactly the case the forge presents.
2. **Extend the existing source override to the names and the certificate, without naming
endpoints.** Rejected: it keys on a port number. A module's route contribution names a port as
well, and nothing says the two are the same thing, so one statement cannot be made to reach all
three mechanisms. Naming the endpoint is what makes that possible.
3. **Derive reach from whether the node has a public domain recorded.** Rejected: that is a property
of the machine, and two endpoints on one machine differ — a database and a web front end on the
same host.
4. **A fourth reach for "public name, internal authority"** — a name that resolves publicly and must
not appear in a public issuance log, obtained by DNS-01. Deferred, not rejected: it is a real case
and it is a question about which challenge an authority uses, not about how far an endpoint
reaches. Left to the certificate work as an open question.
5. **Make the manifest silent on reach and require every assignment to state it.** Rejected for the
transition: every endpoint reachable today would close until an assignment named it, which is a
flag day across the whole catalogue.
## Decision
**A module declares named endpoints.** An endpoint is one port the module serves, with a name the
module chooses, its protocol, and what it is for. A route contribution **names the endpoint it
routes** rather than repeating a port number. The manifest says what the module serves and what it
would serve it to by default; it does not say what this installation does with it.
**An assignment binds each endpoint and says how far it reaches.** Per node: the machine port the
endpoint is published on, and its **reach** — one of `internal`, `public` or `both`. An assignment
that states nothing keeps the manifest's default, so no machine changes until an assignment says so.
**Reach means all three mechanisms at once, and is the only thing that decides them.**
- `internal` — the filter opens the machine port to the private network; the proxy serves the
internal name and not the public one; the certificate comes from the mesh's own authority.
- `public` — the filter opens it to anywhere; the proxy serves the public name; the certificate
comes from the public authority.
- `both` — both names, each from its own authority, and the filter opens to anywhere.
**An endpoint that is not routed is reached but never named.** An endpoint with no route contribution
yields filter rules and nothing else: no name is composed and no certificate is requested. Git over
ssh is that case, and it is the case the model could not express.
**The authority stops being chosen by which names happen to exist.** The proxy composes the names the
assignments asked for, and asks each name's own authority for it. A name nobody asked for is not
composed, so it is not certified.
**The two existing settings are this, completed.** The per-node port mapping becomes the endpoint's
binding. The per-node source override becomes its reach, widened from the filter alone to the names
and the certificate as well.
## Progressive insight — 2026-09-29, from building it
**Reach does not mean the same thing to the filter for an endpoint the proxy serves.** The decision
above says `internal` means "the filter opens the machine port to the private network" and `public`
means "the filter opens it to anywhere". For a routed endpoint the second half is wrong, and
[ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) already said so before this
record was written: *a public service is exposed through the proxy, not by opening its own port* — it
listens `from: mesh`, only the proxy reaches it, and it is exposed by name.
Found by trying to express one real module, not by review. Its routed name must be public, because
browsers post to it; its machine-side port must not be, because that port serves the dashboard in
cleartext. Under one value driving both, saying "public" would have reopened a port an operator had
just closed. Measured the same evening: that module's routed name answered from the internet over TLS
while its machine-side port was refused from the same place. The port is not the path.
So the reach of a **routed** endpoint asks for names, and its port keeps what the manifest said. The
reach of an **unrouted** endpoint — git over ssh, a mail port, the bus — governs the port, because
there is no name and the port is the only way in. That is the same split this record already draws in
*an endpoint that is not routed is reached but never named*; what it got wrong was carrying the filter
across it.
This corrects a fact, not the decision: one statement per endpoint, three things derived from it and
none of them deciding on its own, all stand. The table in the decision should be read with the filter
column applying to an unrouted endpoint.
## Consequences
- **A manifest gains endpoint names, and a route contribution names an endpoint instead of a port.**
Every routed module's manifest changes. The word ships one release before any manifest uses it, and
reaches the build machine and the control plane first.
- **One derived value is read by three things** — the filter's rules, the proxy's contributions, the
certificate request — so they can no longer disagree, and a disagreement becomes a refusal at the
assignment rather than a surprise on a machine.
- **A name that must not be public becomes writable, and therefore checkable.** It also gives
[issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md) a
declared answer to read: which endpoints are internal is what says whose root must be installed
where.
- **[Issue 139](../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md)
becomes answerable**: the endpoint's assignment names the machine that serves it, which is the fact
the internal name should be composed from.
- **Reach becomes reportable.** The mesh can say, per endpoint, where it is reachable from and which
authority holds its certificate — neither of which `status` can say today.
- **This narrows [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md).** Its
decision stands: the firewall is derived and host-applied, not a provider. What no longer holds is
that a listen's `from` is the whole of public-versus-internal; it is the filter's share of a
statement that also governs names and certificates.
- **What got harder:** every endpoint needs a name, including a module that serves exactly one port
and had no reason to name it. And an installation that wants a module public must now say so on the
assignment rather than inheriting it from the definition, which is more to say and the reason it is
right.
## How it is checked
- **One module, two endpoints, different reach.** A module declaring an internal endpoint and a
public one renders a filter opening one to the private network and one to anywhere, asserted per
chain body so a rule in the wrong chain cannot pass.
- **The names follow the reach.** The same module's routed endpoint composes the internal name only
when internal, the public name only when public, and both when both — and a certificate is
requested from the matching authority for each name composed and for no other. This fails against
the previous behaviour, where both names and both certificates are always composed, which is how
it is written.
- **An unrouted endpoint is filtered and never named.** Asserted for an endpoint with reach and no
route contribution: rules rendered, no contribution, no certificate request.
- **An assignment naming an endpoint the module does not declare is refused where it is said**, as is
a reach that is not one of the three — before it reaches a machine, because a ruleset that does not
load is a machine filtering nothing.
- **An assignment that states nothing renders byte-identically to today**, so every machine already
converged is untouched until its assignment says otherwise.
## References
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — narrowed here
- [ADR 0046](0046-a-module-configuration-is-its-assignments-not-its-manifest.md) — where reach belongs
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — the public name this composes
- [ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md) — why reach is not a definition's
- [issue 140](../04-ISSUES/140-an-endpoints-reach-is-not-declared/00-report.md) — the measurement
- [issue 139](../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md),
[issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md)
@@ -0,0 +1,137 @@
---
topic: what runs on it
status: superseded
date: 2026-09-28
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0137-a-machine-says-which-networks-it-routes.md
superseded-by: 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md
---
# 139. A network is forwarded because a module declared it
## Context
[ADR 0137](0137-a-machine-says-which-networks-it-routes.md), decided the same week, gave a machine a
way to say which networks it routes for its guests. It was written because the derived filter's
forward chain allowed two ranges named as constants in the control plane's source — the container
runtime's bridge pool, and part of the pool its compose files are given — with a comment admitting
the gap: *a machine whose runtime is configured with something else needs this to say so, which is a
thing the mesh cannot derive.*
**It can be derived, and from the right place.** Measured on the last machine still to be converged
([issue 141](../04-ISSUES/141-the-forward-chain-does-not-follow-the-modules/00-report.md)): twenty-one
container networks, nine inside the runtime's bridge pool, twelve in the other private range, and six
of those outside the constant's lower bound — so the flip would have cut their guests off exactly as
it did on the workstation that produced 0137.
Naming a range to cover the six is what 0137 provides for, and it is the wrong instrument. Of those
six networks, **four are networks the mesh's own modules declare**, present as network resources in
the node's plan and created by the host because a module asked for them. **Two are the predecessor's
leftovers** — compose networks of services the mesh does not run. Any range wide enough to keep the
four forwards the two as well: a firewall widened by hand to protect networks that should not exist.
The mesh already knows which of the twenty-one are its own, because it made them.
**And the node's configuration is meant to follow the modules assigned to it.** That is the mesh's
founding shape — the machine runs modules, and its files, its filter and its accounts are composed
from what runs there ([ADR 0005](0005-the-node-host.md),
[ADR 0010](0010-delivery.md),
[ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)). The forward chain is the
one derived thing that consults a constant and a list a person types.
## Considered Options
1. **Keep 0137 as it stands** — two constants plus a named list. Rejected: the list is written in
addresses, and addresses are what the runtime allocates, so the only entry safe enough to keep a
machine working is wider than the truth. It cannot distinguish a network the mesh made from one
left behind, which is the distinction that decides whether forwarding it is correct.
2. **Derive it from what the machine reports.** Still rejected, on 0137's own grounds: a test bed
creates its bridge between one declaration and the next, and a filter that follows whatever
appeared on a machine is a firewall that widens itself. **This decision is not that** — see below.
3. **Have the control plane allocate each module network's range from a pool it owns,** so it can
render the address itself. Rejected: more machinery for no gain. The runtime already allocates and
the host already knows, and taking allocation over means the mesh owning an address space it has no
other reason to own.
4. **Have each module declare its network's range.** Rejected by
[ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md): a definition names no address,
and the same definition runs on machines whose runtimes have allocated differently.
## Decision
**A network is forwarded because a module declared it.** Per node, the forward chain forwards the
networks of the modules assigned there, and by default nothing else. A module unassigned stops being
forwarded at the next reconcile.
**The host resolves a declared network to its addresses.** A network resource carries a name; the
runtime allocates the subnet when the network is created. So the control plane declares *forward the
networks these modules asked for* and the host — which made them, and already resolves a container by
its name — renders the addresses. [ADR 0005](0005-the-node-host.md) holds: the host applies, it does
not decide.
**Deriving from the declaration is not deriving from the machine.** Both of 0137's objections fall
away. The set is known before the network exists, because a module declared it, so a network created
between two declarations is already in the one that asked for it. And it cannot widen itself: a
network nobody declared is never forwarded, however it appeared on the machine.
**The runtime's own default bridge is forwarded, from what the runtime reports.** Containers that name
no module network attach to it, and it belongs to the runtime rather than to any module — so the host
renders it from what the runtime says, not from a range named in the control plane. The constants go.
**What a machine says is for guests no module declares.** A test bed is not a module and its range is
not a module's; that is the case 0137's mechanism is for, and it keeps it — added to the derived set,
never replacing it, as 0137 decided. Narrowed to that, it is named for it.
**Their guests keep address and name service**, per declared network, unchanged from
[ADR 0137](0137-a-machine-says-which-networks-it-routes.md): the input chain admits that network's own
DHCP and DNS and nothing else.
## Consequences
- **The two constants are removed**, and with them the class of fault that a machine's guests depend
on a range that describes some other machine.
- **This is a behaviour change, not a refactor.** On the machine measured, the derived set and the
constant do not cover the same ground — that is the whole reason for the record. A machine whose
module networks happen to fall inside the old ranges renders the same rules.
- **A range that exists only to keep a leftover alive becomes visible as such**, because it will not
be in the derived set and has to be said out loud to survive.
- **`node networks` narrows** to guests no module declares, and the preview says which of a machine's
networks are the mesh's and which are not, so the difference is readable before a flip rather than
after.
- **A module's declaration gains nothing.** It already declares its network; what changes is that the
filter reads it.
- **What got harder:** the host renders part of the forward chain from what it created, so the
control plane no longer holds the whole rule set as text. The rule the mesh states is the set of
networks; the addresses are the machine's.
## How it is checked
- **Only declared networks are forwarded.** A node with two modules that declare networks renders
forward rules for exactly those two, and none for a third network present on the machine that no
module declared. This fails against the previous behaviour, which forwards by range and cannot tell
them apart, and that is how it is written.
- **Unassigning a module removes its network's rule** at the next reconcile, asserted on the rendered
chain rather than on the intent.
- **Guests of a declared network keep address and name service**, asserted per chain body so a line in
the wrong chain cannot pass — carried from 0137.
- **The runtime's own default bridge comes from the runtime**, asserted by rendering for a runtime
whose default bridge is somewhere other than the range the constant named.
- **A machine that names a range for guests no module declares still gets it**, added to the derived
set and not replacing it.
- **Each family is matched in its own syntax**, carried from 0137: one set holding both is a syntax
error, and a ruleset that does not load is a machine filtering nothing while its unit reports a
fault.
## References
- [ADR 0137](0137-a-machine-says-which-networks-it-routes.md) — narrowed here; its mechanism keeps the
case it is right for
- [ADR 0005](0005-the-node-host.md) — the host applies; the addresses are the machine's
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the filter is derived from
what runs there
- [ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md) — why a module does not name its
range
- [issue 141](../04-ISSUES/141-the-forward-chain-does-not-follow-the-modules/00-report.md) — the
measurement
- [issue 137](../04-ISSUES/137-converging-a-machine-cut-off-its-own-guests/00-report.md) — the
breakage that produced 0137
@@ -0,0 +1,146 @@
---
topic: what runs on it
status: accepted
date: 2026-09-28
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md
supersedes:
- 02-DECISIONS/0137-a-machine-says-which-networks-it-routes.md
- 02-DECISIONS/0139-a-network-is-forwarded-because-a-module-declared-it.md
---
# 140. The filter constrains what arrives from outside, and says nothing about a machine's own guests
## Context
The filter the mesh derives blocks traffic passing *through* a machine unless something allows it,
because a container's published port is traffic passing through rather than traffic arriving at the
machine itself ([ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md),
[issue 047](../04-ISSUES/047-the-firewall-does-not-cover-published-container-ports/00-report.md)).
Having blocked all of it, the filter then had to let the machine's own containers reach outward again.
It does that by listing the address ranges those containers sit on.
As rendered on a converged workstation today:
```
policy drop
ct state established,related accept
ip saddr 172.16.0.0/12 accept
ip saddr 192.168.128.0/17 accept
ip saddr 10.0.0.0/8 accept
ip saddr 192.168.16.0/20 accept
... four more
```
Two of those ranges were constants in the control plane's source. The rest were typed by the operator
after [ADR 0137](0137-a-machine-says-which-networks-it-routes.md), which existed to make the typing
possible, because converging that workstation had cut every one of its containers off from the
internet and nothing reported a fault
([issue 137](../04-ISSUES/137-converging-a-machine-cut-off-its-own-guests/00-report.md)).
**The list is the mistake, not its contents.** Every attempt to make it correct fails the same way.
A constant describes one machine. A typed range goes stale, and cannot tell a network the mesh made
from one a predecessor left behind — measured on the control-node, where six such ranges fall outside
the constants and two of the six belong to services the mesh does not run
([issue 141](../04-ISSUES/141-the-forward-chain-does-not-follow-the-modules/00-report.md)).
[ADR 0139](0139-a-network-is-forwarded-because-a-module-declared-it.md) tried to generate the same
list from the modules and put half the rule set on the machine to do it. Three records, one list, and
the list should not exist.
**Because the mesh has no policy about a container reaching outward.** What the filter is for is
stated in [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md): which port is open,
and to whom. That is about what arrives. A container of this machine's own opening a connection to
something else is not a port being opened to anybody, and enumerating the addresses it might do so
from is bookkeeping about the machine's internal plumbing, which the mesh neither owns nor can know.
**The system being replaced never had this fault, and its rule says why.** The chain still protecting
the control-node applies only to traffic arriving on that machine's outward link, and leaves
everything else alone. The mesh's filter dropped that distinction and replaced it with a list of
addresses.
## Considered Options
1. **Keep the list and generate it better** — from the modules' declared networks, or from what the
machine reports. Rejected: [ADR 0139](0139-a-network-is-forwarded-because-a-module-declared-it.md)
is that, and it puts part of the rule set on the machine, which makes the rule set partly the
machine's and the derivation advisory.
2. **Name the guest links instead of their addresses, and allow only those.** Rejected as more than is
needed: it fails in the safe direction, but it is still a list that has to keep up with the
machine, and the thing it protects against — a container reaching outward — is not a thing the mesh
has a position on.
3. **Do not block traffic passing through at all.** Rejected: that is
[issue 047](../04-ISSUES/047-the-firewall-does-not-cover-published-container-ports/00-report.md),
where a published port was reachable from anywhere because no rule mentioned it.
4. **Constrain what arrives from outside, and nothing else.** Adopted.
## Decision
**The filter constrains traffic arriving from outside the machine, and says nothing about traffic that
did not.** Traffic passing through the machine is allowed unless it arrived on one of the machine's
outward links, in which case it is allowed only where a declared endpoint's reach admits it
([ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md)). A container of this
machine's own reaching anywhere is not filtered, because the mesh has no position on it.
**A machine says which of its links face outside.** One node-level fact, reported by the machine the
way it already reports the kind of firewall it found and the tunnel it carries — not a setting, not a
list of addresses, and not something anybody types. It does not change when a module is added or
removed, which is what separates it from the list it replaces.
**A machine that has reported no outward link is sent no filter.** Rendering a rule around a link
whose name is not known produces a rule set that does not load, which is a machine filtering nothing
while its unit reports success. The refusal happens in the control plane, where a person reads it, and
the machine keeps the filter it already has.
**No addresses of the machine's own networks appear in the filter.** The two constants are removed and
`node networks` is removed with them, along with everything any machine was told to say through it.
Ports continue to follow the modules exactly as before: a module assigned to a machine opens the port
its assignment says it reaches on, and nothing about a network is said anywhere.
## Consequences
- **Three records collapse into one rule.** 0137 and 0139 are superseded. What 0137 was right about —
that converging a machine had silently cut off its own containers, and that nothing previewed it — is
answered by removing the cause rather than by giving the operator a way to compensate for it.
- **Every machine already converged loses its declared ranges and keeps working**, because the traffic
those ranges allowed is now allowed by not having arrived from outside. The workstation's five ranges
and the laptop's one are deleted rather than migrated.
- **A machine's test beds stop being a special case.** A bed's network is created while the machine
runs and was the case no list could cover; it is now covered by not being mentioned.
- **A new fact travels in the report**, and the control plane refuses to compose a filter without it,
so the order of the roll-out matters: the machines report before the control plane depends on it.
- **A machine with more than one outward link says so**, and a machine that acquires one while the mesh
is not looking is treated as internal until its next report. That window is the cost of this shape;
it is bounded by the report interval, and it exists on machines whose outward link changes, which
are the machines with nothing published to the outside.
- **What got harder:** nothing in the declaration, and one more thing a machine must be able to work
out about itself. A machine that cannot say which link faces outside cannot be given a filter.
## How it is checked
- **A machine's own container reaches outward with no network named anywhere.** A bed converges a
machine carrying containers on several networks, none of them mentioned in any setting, and each
reaches out afterwards. This fails against the previous behaviour, where the same flip cut them off,
and that is how it is written.
- **A port declared reachable from outside is reachable; one that is not, is not.** Probed from off the
machine's private network, for a published port and for an undeclared one, before and after the flip.
- **A network created after the filter was composed needs no new filter.** A network is made on the
machine after its last declaration and a container on it reaches out, with nothing re-sent.
- **No address of a machine's own networks appears in a rendered filter**, asserted on the text so a
range cannot creep back in.
- **A machine that reports no outward link is sent no filter, and the refusal names it** — asserted in
the control plane, and that the machine's existing filter is left alone.
- **A machine reporting two outward links has both constrained**, asserted per chain body so a rule
covering one and not the other cannot pass.
## References
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — what the filter is for
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — what admits traffic
arriving from outside
- [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md) — why traffic passing through is
filtered at all
- [ADR 0137](0137-a-machine-says-which-networks-it-routes.md),
[ADR 0139](0139-a-network-is-forwarded-because-a-module-declared-it.md) — superseded here
- [issue 137](../04-ISSUES/137-converging-a-machine-cut-off-its-own-guests/00-report.md),
[issue 141](../04-ISSUES/141-the-forward-chain-does-not-follow-the-modules/00-report.md)
@@ -0,0 +1,154 @@
---
topic: what runs on it
status: accepted
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0005-the-node-host.md
---
# 141. The host delivers its own successor, and versions live side by side
## Context
[Issue 142](../04-ISSUES/142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md). A
merge builds every changed module and the control plane — which is itself a module — and the result
reaches the machines running it with nobody asking. The host is the exception: it is not a build
target, no declaration delivers it, and every machine in this mesh runs a byte-identical binary that
somebody built on a workstation and copied out.
The half that *recovers* from a bad host exists. `internal/upgrade` can tell that the executable this
process started from was replaced on disk, and it records which version last completed a reconcile.
The launcher counts consecutive failed starts, calls a rollback at the limit, and treats a clean exit
as the host standing aside so that the next loop runs whatever is on disk now. That supervision is
complete and correct.
Two things make it dead code:
- **`Replaced()` is called by nothing but its own tests.** Nothing tells the running host that a
successor is waiting.
- **The rollback resolves a version through the machine's package manager** — `pacman -U` from the
package cache. No machine here has the host installed as a package, so the recovery cannot run on
any of them; and being written in one package manager's terms, it cannot run on two of the three
operating systems the host is built for — [ADR 0005](0005-the-node-host.md) builds one binary per
operating system, pinned at link time.
**The record already points at the answer.** What is kept is a *version*, not a path. Keeping a
version is only useful to something that can choose between versions present on the machine, which is
what the package manager was being asked to do. The versions can simply be on disk.
## Considered Options
1. **Deliver the host as a package, as the rollback assumes.** Rejected: it needs a package built and
a repository trusted per operating system, three of each, and the existing `package` resource
asserts presence and deliberately never a version — "version is the package manager's business and
the mesh does not hold a second opinion about it" — so it cannot ask for a particular host anyway.
Heaviest of the three and the only one that is different on every machine.
2. **Write the new binary over the running one.** Rejected on a fact: a running executable cannot be
truncated, and `archive` opens what it unpacks with `O_TRUNC`. It could be made to write and
rename, which is better hygiene and worth doing for its own sake, but it buys nothing here that
option 3 does not, and it leaves rollback with nowhere to go back to.
3. **Versions side by side; the newest retires the old.** Adopted.
## Decision
**A host version is delivered as an archive into a directory named for it, and never over a running
one.** The declaration names it like any other archive — fetched by digest, the digest checked before
anything is unpacked. Nothing new travels, no new resource kind, and no change to how archives are
applied, because the path being written is not the path being executed.
**The launcher starts the most recently delivered version.** That is what "the newest" means: the
version whose directory arrived last. It reads no pointer and follows no link — the mesh creates no
links ([ADR 0012](0012-the-mesh-creates-no-symlinks.md)) — and the version is in the path, so nothing
has to be told what is running.
**The running host stands aside for a successor, and only between reconciles.** Finding a newer
version delivered, it finishes the reconcile it is in and exits cleanly. The launcher already reads a
clean exit as exactly this and starts what is on disk now. A host that stood aside mid-apply is the
half-configured machine this project exists to prevent, so the check happens at the boundary and
nowhere else.
**A version that completes a reconcile records itself, and retires what came before it.** The
known-good record is written as it is today. Then versions older than the one before the running one
are removed: the running version and its predecessor are kept, which is exactly what a rollback
needs, and nothing else accumulates.
**Rollback starts the previous version instead of reinstalling a package.** At the failure limit the
launcher pins the known-good version and starts that, once. The second failure is still a different
diagnosis — the previously working version does not run either, so it is the machine and not the
binary — and the halt is unchanged. No package manager, no package cache, and the same script on every
operating system.
**A machine says which host version it is running,** on the report it already sends, beside the other
facts it states about itself. Without it nothing can say a machine is behind, so "every machine
current with its source" cannot include the host.
## Progressive insight — 2026-09-29, the same day
**The delivery is not "nothing new", and this record said it was.** The decision above stands and is
built: versions side by side, the newest runs, the running host stands aside between reconciles, a
completed reconcile retires what is older than the predecessor, rollback picks a directory. What was
wrong was a claim about how a version reaches a machine. The paragraph on delivery said the
declaration "names it like any other archive… nothing new travels, no new resource kind"; the second
half is true and the first is not, because two things the delivery needs do not exist:
- **Nothing can compile it.** A `bundle` artifact is compiled by a closed list of toolchains —
typescript and python — whose own comment says adding a language is a decision, because a language
used by *modules* needs an SDK carrying the broker client, the event envelope and tool serving. The
host uses none of that: it is what applies modules, not one of them. So the obligation that list
warns about attaches to a module written in a language, not to the language being buildable, and
the control plane — also written in Go — is built as an image from a Dockerfile rather than through
a toolchain at all.
- **A version cannot reach the path.** An `archive` resource names a fixed path in the manifest, and
nothing interpolates the built version into it, so nothing can ask for
`…/versions/<version>/`.
Neither changes what was decided, which options were weighed, or any consequence: the shape is
unaffected and the host half is merged and tested. What it changes is the cost, which this record
understated as none. The remaining work is a way to build the host and a way to name a version in a
path, and until both exist nothing delivers a version and every machine takes the fallback — which is
what every machine does today.
## Consequences
- **The host becomes a build target and a module** — a module whose resource is the next host, applied
by the host that is running. The bootstrap is not circular because the two are different versions in
different directories.
- **Rollback becomes usable on every machine**, having been usable on none. It also stops being
written in one operating system's terms.
- **One copy by hand remains, once.** The first host that understands versioned directories cannot be
fetched by a host that does not. That copy is the last, and it is the honest cost of the change
rather than a step in the design.
- **Two versions occupy disk instead of one.** About nine megabytes. The predecessor is the price of a
rollback that does not depend on a cache somebody else may clean.
- **What got harder:** a host must now be able to find its own successor and to judge when it is safe
to stand aside. Both are between reconciles, which is the only moment the host is not mid-change.
- **A machine that is never told a newer version keeps running what it has**, indefinitely and
visibly, because its report says which version that is.
## How it is checked
- **A delivered version is run, and the old one is not.** A bed delivers a second version to a machine
running the first; the host exits between reconciles, the launcher starts the new one, and the
machine reports the new version. This fails against the previous behaviour, where nothing notices a
delivered version at all.
- **It stands aside between reconciles and never inside one.** Asserted by delivering a version while
an apply is in flight: the apply completes, and the exit follows it.
- **A version that will not start is rolled back to its predecessor, once**, and the second failure
halts with the machine named rather than the binary — asserted with no package manager involved.
- **A completed reconcile retires what is older than the predecessor**, and never the predecessor
itself, because that is what a rollback needs. Asserted on the directory afterwards.
- **The report names the running version**, asserted end to end rather than on the function that reads
it, since the point is that the control plane can tell a machine is behind.
- **The launcher picks the newest delivered version** with no pointer file and no link, asserted by
delivering two and checking which runs.
## References
- [ADR 0005](0005-the-node-host.md) — the host, and what its supervision is for
- [ADR 0010](0010-delivery.md) — a declaration is owned resources; this adds no kind to it
- [ADR 0012](0012-the-mesh-creates-no-symlinks.md) — why the version is in the path
- [ADR 0005](0005-the-node-host.md), *it is built per operating system* — why a rollback written in
one package manager's terms was wrong for two of three
- [issue 142](../04-ISSUES/142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md) — the
measurement
@@ -0,0 +1,158 @@
---
topic: the mesh
status: accepted
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0141-the-host-delivers-its-own-successor.md
---
# 142. The mesh delivers its own components as binaries, not as container images
## Context
Measured on the control-node, 2026-09-29:
| what | how it runs | publishes |
|---|---|---|
| host | a binary on the machine | — |
| controller, catalogue, builder, vault | containers | nothing |
| store, registry, broker | containers | ports |
**The mesh's own software is delivered two ways, and the difference is not a property of the
software.** The host and the controller are both written in the same language, both the mesh's own,
both doing the mesh's own work. One is an image fetched from a registry. The other is a file somebody
copied to four machines, owned by no package, built by nothing
([issue 142](../04-ISSUES/142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md)).
**The reason is not a judgement about either, it is that images are the only delivery that works.**
There is no way to put a binary on a machine. The host is hand-copied because of that, and the
controller is an image because of that. Neither was chosen on its merits.
What it costs, all of it measured rather than argued:
- **Genesis must raise a container runtime before the control plane can exist.** The bundle carries
three images and one of them is the controller, *"in the bundle for the same reason they are: there
is nothing to fetch it with yet"*
([design 07](../03-DESIGN/01-to-be/07-the-foundation.md)). So the hardest moment in the mesh's life
has a prerequisite that the thing being started does not need.
- **Updating the control plane depends on the control plane.** Its image is fetched from the registry,
which is a container the controller manages.
- **A change to the host cannot be rolled out at all.** Every machine here runs a byte-identical
hand-copied binary. A change merged yesterday reached none of them.
- **Compiling the language the mesh is written in is not a capability of the builder.** The bundle
toolchains are typescript — real, with a registered base module — and python, which is named in the
list and absent from the catalogue. The controller is built as an image from a Dockerfile, which is
the per-repository incantation the bundle toolchain exists to abolish
([design 18](../03-DESIGN/01-to-be/18-building-a-module.md)).
The half that *receives* a binary safely is already built and tested
([ADR 0141](0141-the-host-delivers-its-own-successor.md)): versions side by side in directories named
for them, the newest run, the running one standing aside between reconciles, retirement keeping the
predecessor, and a rollback that chooses a directory. What is missing is everything that puts one
there.
## Considered Options
1. **Leave it as it is.** Rejected: it is not a design, it is the reach of one mechanism. And it is
what makes a host change undeliverable.
2. **Containerise the host too**, so everything is delivered one way. Rejected: the host is what
starts the container runtime and what applies containers. A host in a container is the bootstrap
problem made total, and the machine would have no way back from a bad one.
3. **Deliver the mesh's components as operating-system packages.** Rejected for the reason
[ADR 0141](0141-the-host-delivers-its-own-successor.md) rejected it for the host: a package and a
trusted repository per operating system, three of each, and the `package` resource asserts presence
and deliberately never a version.
4. **Binaries for the mesh's own components, containers for third-party software.** Adopted.
## Decision
**The mesh's own components are delivered as binaries on the machine.** The host, the controller, the
catalogue, the builder, the vault — the software this project writes. They are delivered by the
mechanism [ADR 0141](0141-the-host-delivers-its-own-successor.md) built: an archive, fetched by
digest, unpacked into a directory named for its version, with the running one standing aside between
reconciles and a rollback that chooses the predecessor.
**Third-party software stays a container.** The store, the registry, the broker. They are somebody
else's build, they are already adopted as modules
([ADR 0078](0078-the-store-and-broker-are-modules.md)), and an image is the right way to carry
somebody else's software. **The container runtime remains required** — modules use it — so this
removes a dependency from the control plane, not from the machine.
**The builder compiles the languages the mesh is written in.** A toolchain for Go, with a base module
providing the compiler, exactly as typescript has. The obligation the toolchain list warns about — an
SDK carrying the broker client, the envelope and tool serving — attaches to a *module* written in a
language, not to the language being compilable. None of these components is a module in that sense;
the host is what applies modules.
**An artifact says what it targets.** A compiled binary is per operating system, pinned at link time
([ADR 0005](0005-the-node-host.md)), and a toolchain deliberately takes nothing from the module,
because anything a module could override there it would be writing a Dockerfile to override. So the
target is a property of the artifact rather than of the recipe, and one artifact declared per target
is one build each.
**A component's version comes from where it sits, not from its linker.** It is unpacked into a
directory named for its version, so it can read its own version from its path. The stamp goes, and
with it the need for a build to know what it will be called.
**Genesis carries a binary reference where it carried an image reference.** The principle does not
change — the bundle names a thing by digest and the host fetches it, pinned because nothing can
resolve a version when no mesh exists — and the container runtime stops being a prerequisite for the
control plane. It stays a prerequisite for the store and the broker, which is where it belongs.
**The order is staged, and each step stands alone.** Compiling Go; an artifact naming its target;
delivering a binary; the host as the first component delivered; the controller, catalogue, builder and
vault out of their containers; genesis last. Genesis is last for the reason it is always last: it
matters for a machine nobody has yet, and every earlier step is provable on a mesh that exists.
## Consequences
- **One delivery for the mesh's own software**, so a change to the host ships the way a change to the
controller does, and neither is copied by hand.
- **The control plane stops depending on a container runtime and on its own registry.** Both remain on
the machine for other reasons; neither gates the control plane's own life any more.
- **`Replaced()`, the known-good record and the launcher's rollback stop being dead code.** They were
written for this and have been called by nothing but their tests.
- **Four more components gain a rollback they do not have.** Today a bad controller image is recovered
by an operator; under this it is recovered the way a bad host is.
- **Two versions of each component occupy disk.** Around nine megabytes each. The predecessor is what a
rollback needs.
- **Genesis gets smaller, not larger.** One fewer image to carry and one fewer runtime to raise before
the control plane.
- **This does not make the components smaller or simpler.** They are the same programs; what changes is
how they arrive. A reader expecting the containers to have been hiding complexity will not find any.
- **What got harder:** the builder gains a language, artifacts gain a target, and the mesh gains a
second kind of thing it must deliver correctly — one where getting it wrong takes the control plane
down rather than a module. That is why the host is first: it is the component whose recovery is
already built and tested.
## How it is checked
- **A component is delivered and runs, with nothing copied by hand.** A bed builds the host from its
repository, delivers it to a machine running an older one, and the machine reports the new version.
This fails today at the first step, because nothing builds it.
- **Each target is built once and only the matching one is delivered.** Asserted by declaring an
artifact per operating system and checking that a machine is offered the one it can run — a host
built for another is what ADR 0005's link-time pin exists to refuse.
- **A component reads its version from its path**, asserted by unpacking the same bytes into two
differently named directories and seeing each report its own.
- **A bad component is rolled back without an operator**, for the host first: a version that will not
start is replaced by its predecessor once, and the second failure halts naming the machine.
- **The control plane comes up with no registry reachable**, which is the dependency this removes —
asserted by raising it with the registry stopped.
- **Genesis raises a control plane with no container runtime running**, and raises the store and the
broker afterwards. Last, and on a machine with nothing on it.
- **A published port count that does not change.** The mesh's own components publish nothing today, so
moving them out of containers must not open anything — asserted on the machine's reachable set before
and after, which the converge preview already reads.
## References
- [ADR 0141](0141-the-host-delivers-its-own-successor.md) — the receiving half, already built
- [ADR 0005](0005-the-node-host.md) — the host, its supervision, and one binary per operating system
- [ADR 0078](0078-the-store-and-broker-are-modules.md) — why third-party software stays a container
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — what genesis must raise, and in what order
- [issue 142](../04-ISSUES/142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md) — the
measurement that started this
- [design 07](../03-DESIGN/01-to-be/07-the-foundation.md) — the bundle's three images, one of them the
controller
@@ -0,0 +1,135 @@
---
topic: what runs on it
status: superseded
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0010-delivery.md
superseded-by: 02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md
---
# 143. A consumer verifies the grant it is given
## Context
A **grant** is what the mesh writes on a consumer's machine so it can reach a provider. The real one
the forge receives for its database, as it arrives:
```
provision postgres-database
at <the provider's machine, by name>
port the machine port the provider is published on
as the role the provider created for this consumer
```
with the credential sealed in a separate file. Four facts and a password, and they are the whole
mechanism by which anything in the mesh reaches anything else.
**The mesh asserts that claim and never finds out whether it is true.**
[Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md):
converging a machine dropped the path from a container to a port on its own machine, and for eleven
hours the mesh answered *all doing what they were told, all heard from, every module current with its
source* while a web application logged, six thousand times:
```
connection to server at "<the machine>" (10.10.0.1), port 6852 failed: timeout expired
```
Every check the mesh makes passed, because every check it makes is about the relationship between the
mesh and a machine: the declaration was applied, the digest matched, every container named was running.
None of them asks whether a consumer can reach what it requires — though the mesh composed the grant
and therefore knows the consumer, the machine, the address, the port and the credential.
**And where the check runs decides whether it catches anything.** The rule in force admitted the
machines' own addresses on the private network. A dial from the *machine* to its own address carries
exactly such a source address, so a check run by the host on its own behalf would have matched that rule
and passed — while every container on the machine was refused. This is inference from the rule that was
loaded, not a measurement: the fault was found and fixed before anyone thought to dial from the host.
It is enough to decide the question, because a check whose position differs from the consumer's is
testing something nobody asked about.
## Considered Options
1. **The control plane dials each provision.** Rejected, and it is the tempting one because the control
plane holds every fact. It sits on the provider's machine for most provisions here and reaches the
address by a path no consumer uses; in the measured outage it would have passed throughout.
2. **The host dials on the consumer's behalf, from the machine.** Rejected for the reason above: the
machine's network position is not the consumer's, and the one outage this exists to catch is exactly
a difference between them.
3. **Ask the module.** Rejected: a module is arbitrary software that the mesh does not write. Some could
report on their provisions and most cannot, and a check that covers the modules that opted in tells
nobody anything about the rest.
4. **Read the module's logs.** Rejected: the failure was in a log the whole time, and reading a module's
logs makes the mesh depend on the wording of software it does not control.
5. **The consumer verifies it, from its own network position.** Adopted.
## Decision
**A consumer verifies each grant it is given, from its own network position.** After a reconcile has
applied a grant, the machine opens a connection to the address and port that grant names, from inside
the consumer's own network namespace — the same position the consumer's software dials from, which is
the only position that answers the question the grant asks.
**It is a connection, not a conversation.** Whether the port accepts a connection is what a grant
claims; whether the credential is right, the role exists or the schema is current is the provider's to
answer and the consumer's to discover. A check that spoke each provision's protocol would be a second
implementation of every provision, and would fail for reasons that are not the mesh's.
**One failure is not news.** A provider restarting is ordinary, and so is a consumer between containers.
A grant is reported unreachable only after it has failed on **consecutive** reconciles, and the count is
what the machine reports rather than the last attempt — so a reader can tell "it was briefly away" from
"it has never worked".
**A grant that cannot be checked is said to be unchecked, never assumed good.** A consumer that is not
running has no network position to dial from; that is not a broken grant and must not read as one. It is
also not a verified grant, and the two are different sentences.
**What it costs to be wrong is the constraint on all of it.** A check that reports a working provision
broken trains a reader to ignore the report, which is worse than having none — the fault this
repository keeps finding, one level up. So the threshold is consecutive failures, the check is the
cheapest thing that answers the question, and an unknown is reported as unknown.
**The mesh says it where it says everything else.** A machine's report carries its unreachable grants,
and `status` names them beside what is out of date — so "every module current with its source" stops
being the whole of what the mesh will tell you about a machine whose modules cannot reach each other.
## Consequences
- **The mesh can be wrong out loud.** It has been able to assert a grant and not check it; now a grant
that does not work is a thing the mesh says, and the eleven hours of issue 145 become minutes.
- **The host gains the ability to act from a container's network position**, which it has not needed
before. That is a real capability and the only one this needs.
- **A machine reports something that is not about the declaration.** Everything it reports today is
what it applied and what it holds; this is the first thing it says about whether what it applied
works.
- **A provision with no port is not checked**, because there is nothing to dial. Several are files and
secrets, and saying "checked" about those would be the appearance of verification that this record
exists to remove.
- **What got harder:** a reconcile does more than apply. Every grant adds a connection attempt on a
cadence, which is cheap individually and worth naming: a machine with many consumers dials once per
grant per reconcile.
## How it is checked
- **The outage is caught.** A bed drops the path from a consumer's network position to a provider's
port while leaving the machine's own path to it open — the exact shape of issue 145 — and the grant
reads unreachable. This fails against the previous behaviour, where nothing reported anything, and
against a check run from the machine, which passes while the consumer cannot reach it.
- **A restarting provider is not an outage.** One failed reconcile reports nothing; the count rises and
falls, and the grant reads reachable again without anybody acting.
- **A consumer that is not running reads unchecked, not broken**, asserted separately from the
unreachable case because they are different sentences.
- **A provision with no port is not claimed to be checked.**
- **The report carries the count, not the last attempt**, so "briefly away" and "never worked" are
distinguishable by a reader who sees only the report.
- **`status` names an unreachable grant**, asserted on the output, since a check nothing surfaces is
the same as no check.
## References
- [ADR 0010](0010-delivery.md) — the declaration is owned resources; a grant is one of them
- [ADR 0009](0009-modules-and-the-graph.md) — what a provision and a consumer are
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
— the eleven hours
- [issue 136](../04-ISSUES/136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md) — the
same distance between a declaration and a machine, one level down
@@ -0,0 +1,121 @@
---
topic: what runs on it
status: accepted
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md
supersedes: 02-DECISIONS/0143-a-consumer-verifies-the-grant-it-is-given.md
---
# 144. Anything on a machine may call anything on it, and that is the whole of "local"
## Context
Everything in the mesh should be able to call:
- what runs on the same machine;
- another machine's service over the private network, if that service is exposed there;
- another machine's service over the public network, if it is exposed there.
Three cases. The filter had two of them.
**The first was broken and the break was invisible.** A service exposed to the private network rendered
as the machines' own addresses on it. A caller on the machine carries such an address; a caller inside
one of that machine's containers carries a bridge address and matched nothing. Measured:
```
the machine: local 10.10.0.1 dev lo src 10.10.0.1
a container: 10.10.0.1 via 172.17.0.1 dev eth0 src 172.17.0.8
```
Same destination, same machine, two source addresses. The rule named the first and silently refused the
second, so a module reaching its database on its own machine's name timed out for eleven hours
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)).
**The second case works, and by accident.** A caller on another machine reaches the private network over
the tunnel, and arrives carrying that machine's own address — so the rule matches. It would not have
matched the caller's own address either; the tunnel rewrites it. That two of three cases worked is why
this looked correct.
**[ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) answered the wrong question.** Written
hours earlier, it proposed that a consumer verify each grant it is given by opening a connection from
its own network position — and it went to some length about *which* position, because whether a caller
sat in a container changed the answer. That difference was the bug. A verification mechanism would have
reported this outage sooner and would not have prevented it, and the machinery it needed existed only
because the rule was wrong. The remedy for a configuration error is the correct configuration.
**And a module is not a container.** A module is software that delivers one or more services, and it may
do that as a container, an installed package with a unit, a binary, or files something else reads. Of 72
modules in the catalogue, 61 happen to use a container and 11 do not — among them the resolver, the ssh
daemon and the intrusion-prevention module. A rule that reasons about containers describes most of the
mesh and not the mesh.
## Considered Options
1. **A line per service admitting the machine's own callers.** Rejected: it is what was written first,
and it only ever covers the services somebody remembered to think about. It also states, service by
service, a thing that is true of the machine.
2. **Verify each grant from the consumer's position** ([ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md)).
Rejected as a remedy: it observes the fault rather than removing it, and the question it agonised over
— which network position — exists only while the fault does.
3. **Enumerate the addresses a machine's callers may have.** Rejected for the reason no address is named
anywhere in this filter any more ([ADR 0140](0140-the-filter-constrains-what-arrives-from-outside.md)):
a range describes one machine and goes stale in silence.
4. **Local is not filtered, stated once.** Adopted.
## Decision
**Anything on a machine may call anything on that machine, and the filter says so once.** Not per
service, not per port, and not by naming who the callers are: traffic that did not arrive from outside
the machine and did not arrive over the private network is the machine's own, and is admitted. It is
asked by the link the traffic arrived on, because that is a fact about the machine rather than a list
that describes one.
**Local is not a boundary this mesh draws.** Whether a caller is a container, a unit, or the operator's
shell changes nothing, because the thing being decided is "is this the same machine" and the answer does
not depend on the form the caller takes.
**The other two cases are unchanged and are now legible beside it.** A service exposed to the private
network admits the machines on it; a service exposed publicly admits anything. Three cases, three lines,
and a reader can see all three at once.
**[ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) is superseded and nothing replaces it.**
Whether the mesh should check that a grant works is a real question — it reported this machine healthy
for eleven hours — but it is a question about what the mesh can say, not about what it should do, and it
must stand on its own rather than as the remedy for a rule that was wrong. It is not built.
## Consequences
- **The three things everything should be able to call are three lines**, and the first is one line
rather than one per service, so a service added tomorrow is reachable locally without anybody
remembering to say so.
- **A form of module stops mattering to the filter.** The 11 modules that are not containers were never
affected by this bug and were never the reason it was hard to see; they are the reason the rule should
never have mentioned containers.
- **The mesh still cannot say when a grant stops working.** That is the live gap, recorded in issue 145
and no longer pretending to have an answer.
- **What got harder:** nothing. This removes a line per service and replaces it with one.
## How it is checked
- **A caller on the machine reaches a service on it, in the input chain**, asserted on that chain's own
body — because the forward chain carries the same line in the same words, and an assertion on the
whole rendered file passed with the input chain's copy deleted. That is what
[ADR 0137](0137-a-machine-says-which-networks-it-routes.md)'s tests already say to do.
- **It is one rule, not one per service.** Asserted by rendering two services of different reach and
refusing a per-port local line.
- **The three reaches render as three lines**, asserted together, so the whole of what the filter says
about who may call what is one test.
- **The measured case:** from a container on the machine, a service exposed to the private network on
that machine answers. This is the outage, and it fails against the rule this replaces.
## References
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the filter is the sum
of what its modules listen on
- [ADR 0140](0140-the-filter-constrains-what-arrives-from-outside.md) — why no address is named
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — internal and public,
the other two cases
- [ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) — superseded here
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
@@ -0,0 +1,120 @@
---
topic: what runs on it
status: superseded
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md
superseded-by: 02-DECISIONS/0146-connectivity-is-checked-by-name-per-hosting-form.md
---
# 145. A module checks what the mesh claims is reachable, and it checks itself
## Context
The mesh asserts three things are callable ([ADR 0144](0144-anything-on-a-machine-may-call-anything-on-it.md)):
what runs on the same machine, another machine's service exposed to the private network, and another
machine's service exposed publicly. It has never checked any of them.
[Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md):
the first of the three was broken for eleven hours and the mesh answered *all heard from, every module
current with its source* throughout. Every check it makes is about the relationship between the mesh and
a machine — applied, current, containers running — and none about whether anything can reach anything.
**A first answer was drafted and withdrawn.** [ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md)
put the check inside the host, verifying each grant from the consumer's network position. It was
superseded because the difference it worked so hard to reproduce — whether a caller sat in a container —
was the bug itself. What survives from it is the part that was right: a check run from the wrong place
proves nothing, and the mesh's own reports are not evidence about the network.
**The mesh already has the shape for this and it is a module.** A module can declare a container that
runs on a cadence ([ADR 0053](0053-a-step-that-runs-on-a-schedule.md), and three modules already use
`*/5 * * * *`), can be given the mesh's roster as a rendered fact — every machine's name, address and
this node's own identity, the same mechanism the resolver and the operator's ssh configuration use — and
can emit what it found on the bus. Nothing new is needed to build this except the module.
**What it must not check is the trap.** The obvious probe target is ssh: present on every machine, never
closed by design. Dialling it would have passed throughout the outage, because ssh is admitted
unconditionally and the thing that broke was a service exposed to the private network. A checker whose
probe is unconditionally open measures the one path that cannot fail, which is the failure this whole
sequence keeps producing — a check that reads as verification and verifies nothing.
## Considered Options
1. **The host verifies each grant** ([ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md)).
Superseded. It needed the host to act from another network position, which is machinery that exists
only while local calls are filtered wrongly.
2. **The control plane dials every node.** Rejected: it sits on one machine and reaches the others by a
path no ordinary caller uses. It would have passed throughout the outage.
3. **Probe an existing service.** Rejected for the target problem above: the services guaranteed on every
machine are the ones that are never closed, so they cannot fail the way the mesh fails.
4. **A module on every machine that serves its own probe and dials the others'.** Adopted.
## Decision
**A module runs on every machine, serves an endpoint of its own, and dials every other machine's.** The
probe is the module's own endpoint, declared reachable over the private network — so the thing being
dialled is admitted by exactly the rule that governs every other internally-exposed service, and fails
when that rule is wrong. A second endpoint, declared public, does the same for the public path where a
machine has one.
**It checks the three cases the mesh claims, by name:**
- its **own machine**, by dialling its own machine's address — the case that broke, and the only one that
distinguishes a caller on the machine from a caller in one of its containers;
- **each other machine over the private network**;
- **each machine's public path**, where one is recorded.
**It resolves before it dials, and says which failed.** A name that does not resolve and a port that does
not answer are different faults with different owners, and a checker that reports one sentence for both
sends a reader to the wrong place.
**It runs where the callers run.** The module's own code in its own container, on the cadence the mesh
already has, from the same position as every other module on that machine. It is not the host and not the
control plane, and that is the whole point.
**It says what it found and nothing else.** It emits results; it repairs nothing, opens nothing and holds
no credential beyond its own. A checker that fixes things is a second control plane.
**One failure is not a fault.** A machine rebooting is ordinary. A path is reported broken after it has
failed on consecutive runs, and the count travels with the result so a reader can tell "briefly away"
from "never worked" — the one thing [ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) got
right and worth keeping.
## Consequences
- **The mesh gains the ability to be wrong out loud about the network.** Eleven hours becomes two runs.
- **It is a module, so it is assigned, built, pushed and reported on like everything else** — no new host
capability, no new vocabulary, nothing in the control plane that has to know about checking.
- **Its own endpoint is the instrument.** That is what makes it able to fail; it also means the checker
must be assigned to a machine before that machine can be checked, and a machine without it is
unchecked rather than healthy.
- **It cannot check what it cannot be told.** The roster gives it machines; it does not give it every
module's endpoints, so this checks the paths the mesh claims and not every grant in the mesh. That is
the honest scope of a first one, and the difference is worth saying rather than growing quietly.
- **What got harder:** one more module on every machine, and a module whose whole purpose is to fail
visibly when something else is wrong. Its own failures will be read as the mesh's, which is the cost of
an instrument.
## How it is checked
- **It catches the measured outage.** A bed closes the path from a container to a service exposed to the
private network on its own machine — issue 145's shape — and the checker reports its own machine
unreachable while every other path still reads reachable. This fails against a probe on a port that is
never closed, which is the wrong target this record exists to name.
- **A machine rebooting is not a fault**: one failed run reports nothing, the count rises and falls.
- **A name that does not resolve is reported as that**, not as a port that did not answer.
- **It reports and does not act**: asserted by giving it a broken path and checking nothing on the machine
changed.
- **A machine without the module reads unchecked**, never healthy — asserted on what the mesh says about
a machine it is not assigned to.
## References
- [ADR 0144](0144-anything-on-a-machine-may-call-anything-on-it.md) — the three things that must be callable
- [ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) — superseded; what survives is that the
position matters
- [ADR 0053](0053-a-step-that-runs-on-a-schedule.md) — the cadence
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — internal and public,
which the probe endpoints declare
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
@@ -0,0 +1,125 @@
---
topic: what runs on it
status: accepted
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0145-a-module-checks-what-the-mesh-claims-is-reachable.md
supersedes: 02-DECISIONS/0145-a-module-checks-what-the-mesh-claims-is-reachable.md
---
# 146. Connectivity is checked by name, per hosting form, with a valid certificate
## Context
[ADR 0145](0145-a-module-checks-what-the-mesh-claims-is-reachable.md) decided that a module checks what
the mesh claims is reachable, from where the callers are, because the mesh reported four machines healthy
for eleven hours while a module could not reach its database
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)).
That decision stands. What it got wrong is everything about *what* is dialled.
It dialled a raw port on each machine's address. Three things are wrong with that:
- **A raw port is not how anything in this mesh is reached.** A real caller resolves a name, the proxy
answers it, and the proxy reaches the service. A check that dials a port tests the last hop of a path
with four hops in it, and the three it skips — resolution, the proxy, the certificate — are where most
of the mesh's connectivity actually lives.
- **It tested one hosting form.** A module is software that delivers services, and it may deliver them
from a container, from a unit the mesh writes for its own code, or from a unit a package ships. Those
are three different paths to the same machine, and the outage that produced this was two of them
disagreeing. A probe served one way measures one way.
- **It said nothing about certificates.** An internal name that resolves, routes and answers over TLS
that nothing can verify is not a working path; it is a working path for whoever holds the proxy's
trust and nobody else.
## Decision
**Each hosting form gets its own endpoint, its own route and therefore its own name.** On every machine:
| name | what serves it |
|---|---|
| `connect-docker.<node>.internal` | a container |
| `connect-process.<node>.internal` | the mesh's own code, in a unit the mesh writes |
| `connect-unit.<node>.internal` | a unit a package ships |
and the same set under each machine's public domain where it has one — `connect-docker.<domain>` and its
siblings. The names are the instrument: a failure reads as *`connect-docker.g14.internal` did not answer*,
which says which machine and which hosting form without anybody interpreting anything.
**Every machine checks every machine, by name, over TLS, verifying the certificate.** Not a port, not an
address: resolve the name, connect, complete the handshake, check the certificate against the authority
that should have issued it — the mesh's own for an internal name, a public one for a public name. That is
the whole path a real caller takes, and each step failing is reported as itself.
**No name is written anywhere.** The machines come from the roster the mesh already renders as a fact, and
the labels are the module's. A machine that joins appears in every other machine's roster on the next
push, and they begin checking it without an edit.
**And the module arrives on a machine because the machine exists, not because somebody assigned it.** A
machine that joins and does not have it is worse than unchecked: every other machine is already dialling
its names, so it reads as broken everywhere until someone notices. This is the part the mesh cannot
currently express — see below — and it is the part that makes the rest safe.
**What survives from 0145**, unchanged: it reports and repairs nothing; one failure is not a fault and a
path is broken after consecutive runs with the count travelling with the result; findings are said on the
bus, because a finding in a file on the machine is what this exists to end; and the bus is the one path
that cannot report its own failure, so an emit that does not land is written locally and nowhere else.
## What this needs that the mesh does not have
Named here rather than assumed, because each is a decision of its own and this record is not the place to
make them:
1. **A module that every machine has.** `ScopeNode` means *at most one holder per node* — an exclusivity
rule, not an obligation — and nothing assigns a module at enrolment. Today the resolver, the packet
filter, ssh and intrusion prevention are each assigned per machine by hand, which is the same gap
wearing different clothes.
2. **A container running a module's own bundle.** A `process` runs the mesh's own compiled code with no
image; a `container` needs an image of the module's own, which means a Dockerfile — the thing the
`bundle` artifact exists to abolish. Nothing in the catalogue runs a bundle in a container, so
`connect-docker` has no shape yet.
3. **A unit a package ships, for `connect-unit`.** The `service` resource puts an existing unit into a
state and deliberately installs none, so this form needs a package that serves a port — and naming a
program the machine may not have is
[issue 136](../04-ISSUES/136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md).
4. **A machine's public domain in the roster fact.** The fact carries each machine's name, mesh name,
address and operator account. The public names cannot be composed without the domain.
5. **Something that installs the mesh's own root on a machine.** This is
[issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md), open
since before any of this. Until it is closed, every internal name will fail certificate verification
from every machine — correctly, because nothing can verify it. That is the checker working, and it is
worth saying in advance so the first run is not read as the checker being broken.
## Consequences
- **A failure names the machine and the hosting form.** That is the whole gain over a port: eleven hours
became two runs under 0145, and under this it also becomes one line that says where to look.
- **The checker surfaces issue 129 immediately**, and will report every internal name unverifiable until
it is fixed. A reader must be told that before the first run rather than after.
- **Five things must be built before this is what it says it is**, and until they are, what exists is a
port dial from one position — useful, and not this.
- **What got harder:** a module with three hosting forms of the same trivial service is a strange thing to
read. It is justified only because those three forms are how the mesh actually runs software, and a
checker that tested one of them would keep the class of outage it exists to catch.
## How it is checked
- **A name per hosting form answers from every machine**, asserted by name and not by port.
- **A certificate that does not verify is reported as that**, distinctly from a name that does not resolve
and a port that does not answer — three faults, three owners.
- **A machine that joins is checked by every other machine without an edit**, asserted by adding one to a
bed and looking at what the others dial on their next run.
- **A machine that joins has the module**, which is gap 1 above and is the assertion that cannot be
written yet.
- **The measured outage is still caught**: the path from a container to a service on its own machine is
closed and `connect-docker.<that node>.internal` fails from that machine while the others still pass.
## References
- [ADR 0145](0145-a-module-checks-what-the-mesh-claims-is-reachable.md) — superseded; its core stands
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — the two reaches these
names come from
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — a label plus a domain, which is why no name is written
- [issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md) — what the
internal names will fail on until it is closed
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
@@ -0,0 +1,139 @@
---
topic: what runs on it
status: accepted
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md
---
# 147. A module anchors the mesh's authority on a machine, and takes it away again
## Context
The mesh runs its own certificate authority and every internal name is served with a certificate
from it. No machine trusts it. On an enrolled, adopted workstation — on the private network,
resolving through the mesh's resolver — every internal HTTPS name fails verification with
*unable to get local issuer certificate*
([issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md)).
The certificates are genuine; nothing on the machine has ever been told what issued them.
The authority's only consumer today is a proxy, which fetches the root into a directory of its own
and hands it to one program ([ADR 0098](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md)).
That is enough for the proxy and for nothing else: a browser, `git` over HTTPS, `curl`, a package
manager and every module that calls another module by an internal name read the machine's trust
store, which holds the predecessor's authority and a developer tool's local root, and nothing of
the mesh's.
The predecessor wrote its root into every machine it set up. Removing it was deliberate — an
honest failure beats a name that verifies for the wrong reason — and it leaves the mesh with no
answer at all until this one lands. It is also what keeps the predecessor alive on the machines
that still speak TLS to a mesh name.
**What makes this a decision rather than a patch** is where the knowledge goes. Two mechanisms in
the mesh already write things onto a machine because it is on the private network: `/etc/hosts`
and the registry's plaintext trust ([ADR 0082](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)).
Following that precedent, the controller would inject an anchor into every such machine's
declaration, and issue 129 proposed exactly that. It would work. It would also put *where this
operating system keeps trust anchors* and *which command refreshes its bundles* into the control
plane, for a fact the control plane does not have (the root does not exist until the authority has
run) and a machine that may have no reason to verify a mesh name at all.
## Considered Options
1. **The controller injects the anchor into every machine on the private network**, the
`/etc/hosts` and insecure-registry shape. Rejected: being on the network is what makes the
registry reachable, and that is why network presence is the right trigger *there* — the trust
and the reachability are the same fact. Trusting an authority is not the same fact as being
able to reach it, and the anchor's path and the bundle refresh are a property of the machine's
operating system, which is the host's half of the mesh, not the controller's.
2. **A new host primitive — a `trust-anchor` resource type.** Rejected for now, not on principle.
The host's vocabulary should grow when a shape cannot be said with what exists, and this one
can: a file and a service already express it, as the packet filter proves
([ADR 0140](0140-the-filter-constrains-what-arrives-from-outside.md), whose module writes a
unit file and a service and nothing else). The primitive becomes right the moment a second
operating system is in play, because the anchor directory and the refresh command are exactly
the difference `internal/system` exists to hold. Until then it would be a vocabulary word with
one speaker.
3. **A module that requires the authority, fetches its root, installs it as an anchor and
refreshes the machine's bundles — and removes both when it is no longer assigned.** Adopted.
## Decision
**A machine trusts the mesh's authority because a module put its root there, and stops trusting it
when that module is taken away.**
1. **The module requires `internal-acme-ca`** and reads the provider's bound address and the path
it serves its root at. It requires nothing else and provides nothing: it is a consumer of the
authority like any other.
2. **It fetches the root over the mesh's own network, without prior trust**, because there is no
prior trust to have — this is the module that establishes it — and the network is what
authenticates the fetch ([ADR 0098](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md),
the same reasoning that lets the proxy fetch it). What it accepts is checked: a body that is
not a certificate fails, and the failure is the module's, not a later handshake's.
3. **It installs the root where this machine's TLS clients look, and refreshes the extracted
bundles** — the command that does the refresh is an ordinary part of the unit that places the
anchor, not a new thing the mesh can be asked to do.
4. **Removal is symmetric and is the same unit's business.** Undeclared, the host stops the unit;
stopping it removes the anchor and refreshes the bundles again. A machine that leaves the mesh
stops trusting the mesh, without anybody remembering to go and look.
5. **It is an ordinary assignment.** No machine is given it automatically. A machine that verifies
a mesh name is assigned it, and a machine that does not is not — which is the same statement
the mesh already makes about every other module, and is why this is not the controller's
business.
**One operating system, said out loud.** The anchor directory and the refresh command in the
module today are Arch's. On a machine that is not Arch the unit fails, visibly, rather than
writing a file nothing reads. That is the accurate failure, and it is the signal that option 2
above has become right.
## How this is checked
- **The verification that could not succeed before.** On a machine holding the module, a plain
client fetches an internal HTTPS name with no `-k` and no bundle argument and verifies. On a
machine without it, the same fetch fails with *unable to get local issuer certificate*. Both
halves, because only the pair distinguishes "the anchor works" from "something else already
trusted it".
- **The removal half, in the same bed:** unassign the module, refetch, and the failure returns.
Checking only the arrival is how a trust store fills up with authorities nobody can account for.
- **What is deliberately not checked here:** that the authority issues, that a name resolves, that
the proxy serves. Those have their own beds, and this module's bed passing for those reasons is
the failure mode this record is most exposed to — which is why the negative half is not optional.
**What this bed is dialled at, and why it is the authority itself.** The authority serves its own
API with a certificate it issued, so the handshake under test needs nothing else in the mesh to be
right. A trust bed that reached for a routed name through the proxy would be passing or failing for
the proxy's reasons and the resolver's.
**Written, and not yet run** *(2026-09-29)*. The bed is `trust-anchor` in the lab, and it cannot
execute: raising a foundation fails before any module is reached, in both bundles that exist
([issue 146](../04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md)).
So what stands behind this record today is the rendering — the script the machine would run names
the authority it was bound to, checked in the control plane's own test suite — and **not** a machine
that verified anything. That is a weaker thing than the paragraph above describes, and it stays
written this way until the bed runs.
## Consequences
The predecessor's authority can be retired from a machine once this module is assigned to it,
which is the first time that has been true. `git` over HTTPS to the mesh's forge starts working,
so the ssh-only clone URL stops being a rule. A module on any machine can call another module's
internal name and verify it.
What got harder: one more module to assign to a machine that needs it, and the machine's trust
store now changes when an assignment changes — which is the point, and is also a thing an operator
can be surprised by. The fetch without prior trust is the same exposure ADR 0098 accepted, now on
every machine that holds the module rather than only where a proxy runs: anything that can stand
in the middle of the mesh's own network at the moment of the fetch can be believed. The mesh
already treats that network as the thing it authenticates.
## References
- [issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md) —
the symptom and the evidence.
- [ADR 0098](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md) — a fact made at
first start is fetched from its provider; this extends it from one program to the machine.
- [ADR 0082](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md) — the precedent
this deliberately does not follow, and why it is right where it is.
- [ADR 0005](0005-the-node-host.md) — the host is where one operating system's difference lives.
- [`03-DESIGN/01-to-be/08-connectivity.md`](../03-DESIGN/01-to-be/08-connectivity.md).
+20 -4
View File
@@ -140,6 +140,9 @@ python3 00-META/checks/index.py fail if stale
- **0129** — [A seat carries the protocol of its role](0129-a-seat-carries-the-protocol-of-its-role.md)
- **0130** — [The predecessor is ending, and its broker goes with it](0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md)
- **0131** — [Everything on the mesh speaks to the broker seat, and AMQP is not a provision](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)
- **0132** — [A seat carries the tools its holder must serve](0132-a-seat-carries-the-tools-its-holder-must-serve.md)
- **0134** — [The mesh says what it applied](0134-the-mesh-says-what-it-applied.md)
- **0142** — [The mesh delivers its own components as binaries, not as container images](0142-the-mesh-delivers-its-own-components-as-binaries.md)
### Its tiers, from the bottom up
@@ -203,15 +206,28 @@ python3 00-META/checks/index.py fail if stale
- **0091** — [A mount is declared, and there are three things it can be](0091-a-mount-is-declared-three-ways.md)
- **0099** — [A step that runs once names what it reads, and runs again when it changed](0099-a-step-that-runs-once-names-what-it-reads.md)
- **0110** — [A seat is held by one assignment, from a closed set, and it may deliver a provision](0110-a-seat-is-a-module-assignment-from-a-closed-set.md)
- **0112** — [A module definition names no node, no mesh and no path: everything it needs is a requirement the mesh resolves](0112-a-module-definition-names-no-node-mesh-or-path.md) *(proposed)*
- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md) *(proposed)*
- **0112** — [A module definition names no node, no mesh and no path: everything it needs is a requirement the mesh resolves](0112-a-module-definition-names-no-node-mesh-or-path.md)
- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md)
- **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md) *(proposed)*
- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md) *(proposed)*
- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md)
- **0117** — [A machine's uplink is a seat: the mesh configures the manager, never the link](0117-a-machines-uplink-is-a-seat.md)
- **0118** — [Undeclaring removes what the mesh made, and gives a unit back the state it was found in](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md)
- **0120** — [A roster fact carries its format as a template: the mesh owns the data, the module owns the format](0120-a-roster-fact-carries-its-format-as-a-template.md)
- **0121** — [A system seat is named for its scope, and a module may define its own](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md)
- **0122** — [A seat is data the controller owns, and a rename is a database update](0122-a-seat-is-data-a-rename-is-a-database-update.md)
- **0133** — [A module owns its migrations, and the mesh owns when they run](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md) *(superseded)*
- **0135** — [A module version prepares its state before it runs](0135-a-module-version-prepares-its-state-before-it-runs.md)
- **0136** — [A step gates its module, not the machine](0136-a-step-gates-its-module-not-the-machine.md)
- **0137** — [A machine says which networks it routes](0137-a-machine-says-which-networks-it-routes.md) *(superseded)*
- **0138** — [An assignment binds an endpoint and says how far it reaches](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md)
- **0139** — [A network is forwarded because a module declared it](0139-a-network-is-forwarded-because-a-module-declared-it.md) *(superseded)*
- **0140** — [The filter constrains what arrives from outside, and says nothing about a machine's own guests](0140-the-filter-constrains-what-arrives-from-outside.md)
- **0141** — [The host delivers its own successor, and versions live side by side](0141-the-host-delivers-its-own-successor.md)
- **0143** — [A consumer verifies the grant it is given](0143-a-consumer-verifies-the-grant-it-is-given.md) *(superseded)*
- **0144** — [Anything on a machine may call anything on it, and that is the whole of "local"](0144-anything-on-a-machine-may-call-anything-on-it.md)
- **0145** — [A module checks what the mesh claims is reachable, and it checks itself](0145-a-module-checks-what-the-mesh-claims-is-reachable.md) *(superseded)*
- **0146** — [Connectivity is checked by name, per hosting form, with a valid certificate](0146-connectivity-is-checked-by-name-per-hosting-form.md)
- **0147** — [A module anchors the mesh's authority on a machine, and takes it away again](0147-a-module-anchors-the-meshs-authority.md)
### How it is built
@@ -221,7 +237,7 @@ python3 00-META/checks/index.py fail if stale
- **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md)
- **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md)
- **0016** — [The lab](0016-the-lab.md)
- **0037** — [Where a module lives](0037-where-a-module-lives.md) *(proposed)*
- **0037** — [Where a module lives](0037-where-a-module-lives.md)
- **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md)
- **0068** — [The lab takes requests, one at a time, and runs each from its own copy](0068-the-lab-takes-requests.md) *(proposed)*
- **0069** — [A module is a repository and a path within it](0069-a-module-is-a-repository-and-a-path.md)
+44 -1
View File
@@ -2,8 +2,9 @@
layer: to-be
status: in-progress
code: [mesh-host]
updated: 2026-09-22
updated: 2026-09-29
decisions:
- 02-DECISIONS/0141-the-host-delivers-its-own-successor.md
- 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
- 02-DECISIONS/0102-the-mesh-writes-into-a-shared-file-never-over-it.md
- 02-DECISIONS/0103-what-an-adopted-node-holds-and-what-its-guard-refuses.md
@@ -423,3 +424,45 @@ ignored an instruction and "applied" would be a lie. Applying stays one at a tim
is not. **Checked** by the link's unit tests on the drain, and by the genesis bed's settle wait,
which counts on a node catching up to the newest declaration rather than the oldest.
## The host delivers its own successor
*2026-09-29, from a change to the host that could reach no machine —
[issue 142](../../04-ISSUES/142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md),
settled by [ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md).*
A merge builds every changed module and the control plane, and the result reaches the machines running
it with nobody asking. The host was the exception: not a build target, named by no declaration, and
identical on every machine because somebody had copied it there.
The supervision needed for this was already right. A clean exit from the host means it has stood aside,
and the launcher's next turn runs whatever is on disk. Consecutive failed starts are counted, a
rollback happens at the limit, and a second failure halts with the machine named rather than the binary.
What was missing was smaller than it looked: nothing told the running host a successor was waiting, and
the rollback resolved its known-good *version* through one operating system's package manager, which no
machine here used.
Keeping a version rather than a path was the clue. That is only useful to something that can choose
between versions present on the machine — so the versions live side by side:
- **A version arrives as an archive, in a directory named for it.** The ordinary resource, fetched by
digest and checked before anything is unpacked. The path written is never the path being executed, so
replacing a running binary — which the kernel refuses — never comes up.
- **The launcher starts the most recently delivered version**, reading no pointer and following no
link, because the version is in the path.
- **The running host stands aside between reconciles and never inside one.** Standing aside mid-apply is
the half-configured machine this document exists to prevent.
- **A version that completes a reconcile records itself and retires what is older than its
predecessor.** The predecessor stays, because that is what a rollback needs.
- **Rollback starts that predecessor** instead of reinstalling a package: no package manager, no cache
somebody else may clean, and the same script on every operating system.
- **A machine says which host version it runs**, on the report it already sends, so being behind is
answerable at all.
One copy by hand remains, once: the first host that understands versioned directories cannot be fetched
by a host that does not.
*How it is checked* is stated with the decision — a second version delivered to a running machine is
run and reported; the exit follows an in-flight apply rather than interrupting it; a version that will
not start is rolled back once and the second failure halts; a completed reconcile retires what is older
than the predecessor and never the predecessor; and the newest of two delivered versions is the one
that runs.
+23 -2
View File
@@ -11,8 +11,9 @@ code:
- mesh-catalog modules/postgres
- mesh-catalog modules/lavinmq
- mesh-lab test/integration/mesh.test.ts (a bare machine becomes a mesh)
updated: 2026-09-22
updated: 2026-09-29
decisions:
- 02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md
- 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
- 02-DECISIONS/0088-the-foundation-filters-before-anything-listens.md
- 02-DECISIONS/0004-a-node-and-how-it-joins.md
@@ -314,4 +315,24 @@ of a database and pushed to over the broker. What arrived and what did not is th
**One fault, and it was in the joining.** The token did not say what the mesh calls the machine,
so enrolment needed a flag its own help said it did not — and failed at the broker with an empty
username. Recorded in ADR 0004 as the fifth thing a token carries.
username. Recorded in ADR 0004 as the fifth thing a token carries.
## The mesh's own components arrive as binaries
*2026-09-29 —
[ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md).*
The bundle carries three images and one of them is the controller, *because there is nothing to fetch
it with yet*. That reasoning holds and its conclusion changes: the controller is carried as a **binary**
reference rather than an image reference, pinned by digest exactly as before. Nothing about the bundle's
shape moves — it names a thing and the host fetches it — and the container runtime stops being something
genesis must raise before the control plane can exist. It still raises one, for the store and the broker,
which is where somebody else's software belongs.
The mesh's own components — the host, the controller, the catalogue, the builder, the vault — are
delivered as binaries into directories named for their versions, by the mechanism
[ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) describes. Third-party
software stays a container. The split is not about isolation; it is about who built the thing.
Measured before deciding it: the mesh's own components publish no ports at all, so this opens nothing.
Only the store, the registry and the broker publish, and they are staying as they are.
+118 -1
View File
@@ -7,8 +7,11 @@ code:
- mesh-controller internal/identity/authority.go
- mesh-host internal/identity/serving.go
- mesh-host internal/apply (the service that reflects a rule set)
updated: 2026-09-27
updated: 2026-09-29
decisions:
- 02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md
- 02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md
- 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md
- 02-DECISIONS/0104-a-provision-may-be-answered-by-an-adapter-to-the-predecessor.md
- 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md
@@ -625,6 +628,50 @@ the found firewall reloads and reachable from a container on the node, that a ma
enrols through the openings before and after a reload and a reboot, and that after the flip the
declared port is open and the undeclared one closed.
### It filters what arrives from outside, and not what the machine's own guests send
*2026-09-28, preparing the control-node's convergence —
[issue 141](../../04-ISSUES/141-the-forward-chain-does-not-follow-the-modules/00-report.md), settled by
[ADR 0140](../../02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md), which replaces
[ADR 0137](../../02-DECISIONS/0137-a-machine-says-which-networks-it-routes.md) and
[ADR 0139](../../02-DECISIONS/0139-a-network-is-forwarded-because-a-module-declared-it.md).*
Traffic passing through a machine is filtered, because a container's published port is traffic passing
through rather than traffic arriving at the machine itself. Having blocked it, the filter then had to
let the machine's own containers reach outward again — and it did that by listing the address ranges
they sit on. Two of those ranges were constants in this repository's code, and the rest were typed by an
operator after the flip had already cut a workstation's containers off from everything.
**The list was the mistake, not its contents.** A constant describes one machine. A typed range goes
stale and cannot tell a network the mesh made from one a predecessor left behind — on the control-node,
six ranges fall outside the constants and two of the six belong to services the mesh does not run. The
attempt to generate the list from the modules put half the rule set on the machine and made the
derivation advisory. Three records, one list.
**And the mesh has no position on a container reaching outward.** §4 exists to say which port is open
and to whom, which is about what arrives. A container of this machine's own opening a connection
somewhere is not a port opened to anybody, and the addresses it might do that from are the machine's
internal plumbing, which the mesh neither owns nor can know.
So the filter constrains what arrives from **outside** the machine and says nothing about what did not.
Traffic passing through is allowed unless it came in on one of the machine's outward links, and then
only where a declared endpoint's reach admits it (§6). The machine says which of its links face
outside — one fact it reports, like the kind of firewall it found and the tunnel it carries, not a
setting and not a list of addresses. It does not change when a module is added or removed, which is the
whole difference from what it replaces. A machine that has reported no outward link is sent no filter
at all, and keeps the one it has, because a rule written around a link with no name is a rule set that
does not load — a machine filtering nothing while its unit reports success.
Ports go on following the modules exactly as before: assign a module to a machine and the port its
assignment says it reaches on opens. Nothing about a network is said anywhere, by anybody.
*How it is checked:* a bed converges a machine carrying containers on several networks, none of them
named in any setting, and each reaches outward afterwards — which fails against the previous behaviour,
where the same flip cut them off, and is how it was written; a network made *after* the last declaration
needs no new filter; a declared port is reachable from off the private network and an undeclared one is
not; no address of a machine's own networks appears in a rendered filter, asserted on the text; and a
machine reporting no outward link is refused in the control plane with its existing filter left alone.
## 5 — Certificates
**Two authorities, kept separate on purpose.**
@@ -664,6 +711,22 @@ step, so when the authority moves the root is fetched again and the proxy is rec
*How it is checked:* the route-forwarding bed installs the authority, the proxy and a consumer
from the catalogue and asserts the routed name is served.
**And a machine trusts that authority because a module put its root in its trust store**
([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md)). The proxy's fetch
answers for the proxy and for nothing else: a browser, `git` over HTTPS, a package manager and
every module calling another by an internal name read the machine's own trust store, and the mesh
had never written anything there. A module requiring the authority does the whole of it — fetch
the root over the mesh network, place it where this machine's TLS clients look, refresh the
extracted bundles — and stopping it, which is what being unassigned does, takes the anchor away
and refreshes them again. Not the controller's business, because being on the private network is
what makes the authority *reachable* and is not the same fact as having a reason to *verify* a
mesh name; and because where anchors live and which command refreshes them is one operating
system's difference, which is the host's half of the mesh
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)).
*How it is checked:* on a machine holding the module a plain client verifies an internal HTTPS
name with no bundle argument, and on one without it the same fetch fails to find an issuer — both
halves, because only the pair tells the anchor apart from something that already trusted it.
### What was built
*2026-08-31.*
@@ -710,6 +773,60 @@ that verifies against the internal root and nothing else — which cannot succee
first reached the name to certify it*
([ADR 0066](../../02-DECISIONS/0066-public-routing-is-name-agnostic.md)).
## 6 — One statement behind exposure, filtering and certificates
*2026-09-28, preparing the control-node's convergence —
[issue 140](../../04-ISSUES/140-an-endpoints-reach-is-not-declared/00-report.md), settled by
[ADR 0138](../../02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md).*
The three sections above each decide, independently, how far a service reaches. §3 composes a name
from a label and the node's domain. §4 opens a port to the source a listen named. §5 certifies the
names that exist, from whichever authority the proxy holds. Each is coherent on its own, and together
they mean **reachability is never stated anywhere** — it is the sum of three derivations, and a sum is
not something anyone can review or refuse.
What that costs, measured: an identity provider holding a public certificate valid 90 days and an
internal one valid 24 hours, neither asked for by any assignment, because both names existed and a
proxy certifies what it serves. And an endpoint that is not routed — git over ssh — which can be
spoken about only in the filter's vocabulary, so *this must be reachable from outside* is a setting
exactly one mechanism reads.
**An endpoint is the thing that was missing.** A module declares named endpoints: one port it serves,
what it is for, and what it would serve that to absent any instruction. A route contribution names an
endpoint rather than repeating a port number. An assignment — which is where a module's configuration
lives ([ADR 0046](../../02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md))
— then binds each endpoint to a machine port and says how far it reaches.
One value, three readers:
| reach | the filter opens | the proxy serves | the certificate comes from |
|---|---|---|---|
| `internal` | the machine port, to the private network | the internal name | the mesh's own authority |
| `public` | the machine port, to anywhere | the public name | the public authority |
| `both` | the machine port, to anywhere | both names | each name's own authority |
**An endpoint that is not routed is reached and never named.** No route contribution means no name is
composed and no certificate requested, while the filter still acts on it. That is the case the model
could not express at all, and it is the ordinary case for anything that is not HTTP.
**Nothing moves until an assignment says so.** An endpoint whose assignment is silent keeps the
default its manifest states, so every machine already converged renders exactly as it does today —
the same property §4 needed when a machine gained a way to say which networks it routes.
This is what the certificate questions were waiting for. Which authority signs a name, whether a name
may appear in a public issuance log, and what must be trusted where are all answerable once an
endpoint says whether it is internal — and unanswerable while the proxy decides by composing every
name it can. It is also the fact
[issue 139](../../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md)
needs: an internal name should be composed from the machine serving the endpoint, which is the
assignment that bound it.
**How it is checked** is stated with the decision: one module with two endpoints of differing reach
asserted per chain body; the names and the certificate requests following the reach and failing
against today's behaviour, where both are always composed; an unrouted endpoint filtered and never
named; an assignment naming an endpoint the module does not declare refused where it is said; and a
silent assignment rendering byte-identically to today.
## What this removes
The list is worth having in one place, because it is most of the argument:
+32 -1
View File
@@ -5,7 +5,7 @@ code:
- mesh-controller internal/builder
- mesh-controller cmd/mesh-controller (build, build --behind, push, status)
- mesh-controller internal/inventory/builds.go
updated: 2026-09-21
updated: 2026-09-29
decisions:
- 02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md
- 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md
@@ -226,3 +226,34 @@ when the current failure began and how many reports in a row have said it — th
id, whatever the words; three make the machine stuck, and `status` says so beside the failure. The host keeps trying — stuck is what the mesh
knows, not what the machine is told. *How it is checked:* an inventory test counts three identical
reports, a different one, and a clean apply; the status test asserts the word appears.
## Everything may call what is exposed to it, and local is not a boundary
*2026-09-29, from an outage that ran eleven hours —
[issue 145](../../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
settled by [ADR 0144](../../02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md).*
A grant is four facts and a credential: the provision, the machine, the port, and who the consumer is
when it connects. It is the whole mechanism by which anything in the mesh reaches anything else, and it
rests on three things being callable — what runs on the same machine, another machine's service over the
private network where it is exposed there, and another machine's service over the public network where it
is exposed there.
The filter had two of those. A service exposed to the private network admitted the machines' own addresses
on it; a caller on the machine carries such an address, and a caller inside one of that machine's
containers carries a bridge address and matched nothing. Measured, same destination and same machine:
`src 10.10.0.1` from the machine, `src 172.17.0.8` from a container on it. So a module reaching its
database on its own machine's name timed out for eleven hours while the mesh called the machine healthy.
The second case worked by accident: a caller on another machine arrives over the tunnel carrying that
machine's address, which the rule matched. Two of three working is why this read as correct.
**So local is not a boundary this mesh draws, and the filter says so once.** Traffic that did not arrive
from outside the machine and did not arrive over the private network is the machine's own, and is
admitted — for every service there, not per service. Whether the caller is a container, a unit or a shell
decides nothing, because the question is "is this the same machine".
A verification mechanism was drafted for this and withdrawn. It would have reported the outage sooner and
would not have prevented it, and the part of it that was hard — deciding which network position to check
from — existed only because the rule was wrong. Whether the mesh should check that a grant works is still
open, in issue 145; it is not the remedy for a configuration error.
+26 -1
View File
@@ -5,8 +5,9 @@ code:
- mesh-controller cmd/mesh-builder
- mesh-controller internal/builder
- mesh-catalog modules/builder
updated: 2026-09-25
updated: 2026-09-29
decisions:
- 02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md
- 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md
- 02-DECISIONS/0097-a-vendor-image-is-a-declared-build-input.md
- 02-DECISIONS/0096-an-upstream-image-is-copied-between-registries.md
@@ -263,3 +264,27 @@ ships one and wrong for code the mesh built, which has no unit until the mesh wr
**Tools, hooks and consumers are not further modes**, which is the test of whether three is the
right number: they are loaded by a tool host, and a tool host is a process that stays up.
## The builder compiles the languages the mesh is written in
*2026-09-29 —
[ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md).*
The toolchain list was typescript and python, and only typescript had a base module in the catalogue.
Meanwhile the control plane — written in the language this project is mostly written in — was built as
an image from a hand-written Dockerfile, which is the per-repository incantation this whole mechanism
exists to abolish.
So the list gains Go, with a base module providing the compiler exactly as typescript has one. The
obligation the list's own comment warns about — an SDK carrying the broker client, the event envelope
and tool serving — attaches to a **module** written in a language, not to the language being
compilable. The mesh's own components are not modules in that sense; the host is what applies modules.
**And an artifact says what it targets.** A compiled binary is per operating system, pinned at link
time, and a toolchain deliberately accepts nothing from the module — anything a module could override
there it would be writing a Dockerfile to override. The target is therefore a property of the artifact,
not of the recipe: one artifact declared per target, one build each.
A component's version stops being stamped in at link time. It is unpacked into a directory named for
its version, so it reads its version from its own path, and a build no longer has to know what it will
be called.
@@ -1,17 +1,29 @@
---
layer: to-be
status: proposed
code: []
updated: 2026-09-27
status: in-progress
code:
- mesh-controller internal/catalogue/declaration.go
- mesh-controller internal/catalogue/manifest.go
- mesh-controller internal/link/serve.go
- mesh-controller internal/link/bus.go
- mesh-controller internal/broker/nats.go
- mesh-controller internal/inventory/nodes.go
- mesh-host internal/apply/apply.go
- mesh-tools src/main.ts
- mesh-catalog modules/mesh-catalog
updated: 2026-09-28
decisions:
- 02-DECISIONS/0126-a-module-declares-its-own-seats.md
- 02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md
- 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
- 02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md
- 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0041-events-are-a-relationship.md
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
- 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
- 02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md
- 02-DECISIONS/0135-a-module-version-prepares-its-state-before-it-runs.md
- 02-DECISIONS/0136-a-step-gates-its-module-not-the-machine.md
- 02-DECISIONS/0134-the-mesh-says-what-it-applied.md
---
# 32. What a module declares, and what the bus makes of it
@@ -258,10 +270,30 @@ queue.
and publishes it last-per-subject. A node that was away gets exactly the current one, never a
queue of superseded ones, and a replayed older one is refused by sequence.
**A version prepares its state before it runs.** *Built 2026-09-28.* A module version may declare an
entrypoint that brings its state to the shape that version needs — the same vocabulary as the entrypoints it declares for its
tools and its provisioner, and nothing about how a machine runs it. The mesh runs that entrypoint as it
runs the module's own code, to completion, in the module's own context, and a version whose preparation
did not succeed does not run: the step gates that module and nothing else on the machine
([ADR 0136](../../02-DECISIONS/0136-a-step-gates-its-module-not-the-machine.md)), and the rollout stops
at the first machine that did not take it
([ADR 0135](../../02-DECISIONS/0135-a-module-version-prepares-its-state-before-it-runs.md), superseding
[ADR 0133](../../02-DECISIONS/0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md)).
Once per state, and the mesh derives what a state is: a consumer is a module on a machine, so what the
mesh provisions is per consumer and preparation is too. No level to choose, and no race to lock against.
**Applying is reported to a role.** The host applies and reports to the `mesh-controller` seat —
not to an address it was given at genesis. Held and retried while the store restarts
([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)).
**And the mesh says what it applied.** *Built 2026-09-28.* A report is control traffic only the control plane reads, so the
chain above went dark at the moment it touched a machine: nothing said which version a machine now runs,
or that it refused to. The control plane states those as facts under its own seat's namespace, when what
a machine runs changes rather than on every convergence pass, and anything that cares subscribes the way
the catalogue subscribes to `built` ([ADR 0134](../../02-DECISIONS/0134-the-mesh-says-what-it-applied.md)).
The facts are second-hand by design — one emitter, one ordering — and a machine that cannot reach the bus
produces none, so absence is not health.
What disappears across that chain is every address. No webhook URL, no registered callback, no
"which node is the builder on", no controller endpoint baked into a joining node. That is the
class of bug
@@ -439,6 +471,10 @@ it is the residue of a question the rest of §8 answers and the part a fingerpri
**Whether a module may declare a seat it does not itself claim** — the contract as one thing, the
implementation as another, which is how two competing implementations would ever exist.
**Whether a container should have a readiness notion.** Only an action carries `verify`, so a step that
must run once a service *answers* — seeding through its own API — cannot be declared at all. Named here
because the steps above make the gap obvious, not because they caused it.
**Whether `consumes` naming another module couples too tightly.** It is kept here deliberately —
an event's provenance is its meaning — but a consumer of `billing.order.placed` does depend on
billing existing under that name.
@@ -447,6 +483,11 @@ billing existing under that name.
- **A manifest holds no subject.** A catalogue test: no manifest contains a string matching the
subject grammar. The rule is worthless if it is followed by convention.
- **A preparation is given what the module is given.** A composition test: what the preparation
entrypoint receives equals what the module's own code receives, asserted rather than written twice —
which is the drift a hand-written step invites, three times over in the catalogue today.
- **A convergence that changed nothing says nothing.** Two identical reports, one emitted fact: what is
guarded against is a fact per minute per machine, which is a stream nobody reads.
- **Permissions are exactly the three namespaces.** A composition test per module: the derived
permission set equals what its declaration implies, and a hand-written addition to it fails.
- **A sender cannot read the queue it writes to.** A bed: a module declaring `uses` is refused
@@ -0,0 +1,137 @@
---
layer: to-be
status: designed
code: []
updated: 2026-09-28
decisions:
- 02-DECISIONS/0132-a-seat-carries-the-tools-its-holder-must-serve.md
- 02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md
- 02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md
---
# 33 — The tools the mesh answers
**An agent can call the mesh's tools and cannot find out what they are.** Both halves were measured
on the live mesh on 2026-09-28: a client holding an operator credential connected to the bus, asked a
module for its repositories and got them; the same client's request for the tool list found nothing
serving it. The transport works, the account model works, the adapter that speaks the agent protocol
works. What is missing is the mesh being able to say what it can do.
This design is the answer to that question, and it has three families in it, because a tool belongs to
whoever is accountable for answering it.
## 1. Three families, and why the split is not arbitrary
| Family | Addressed to | Where the definition lives | Example |
|---|---|---|---|
| A **role's** tools | the seat: `mesh.seat.<seat>.tool.<verb>` | the seat's protocol, in the mesh's records | ask *the forge* to list its repositories |
| A **module's** tools | the module: `mesh.mod.<module>.tool.<name>` | that module's code | ask *this gitea* for `gitea_list_repos` |
| The **mesh's** own verbs | the `mesh-controller` seat | the seat's protocol, as above | `status`, `push`, `build`, `assign` |
The split follows accountability. A role is something the mesh guarantees exactly one holder of, so
what the role answers is the mesh's to define and a holder's to implement
([ADR 0132](../../02-DECISIONS/0132-a-seat-carries-the-tools-its-holder-must-serve.md)). A module's
own tools are nobody's business but the module's, and their definitions live where they are
implemented, because a copy kept anywhere else drifts from the code that answers.
The mesh's own verbs are the third family only in where they come from, not in kind: the control plane
holds a seat like anything else, and its tools are that seat's. This is what keeps them addressable
while the control plane is being replaced, which is the moment they are most needed.
**Both names for one capability is deliberate and bounded to this.** A forge holding the `git` seat
answers the role's `list_repos` and its own `gitea_list_repos`, because the same module may run
without the seat — a second instance, kept for one purpose — and then only the second name is true.
The caller chooses which question it is asking. Nothing else in the mesh gets two names.
## 2. What a seat's tool is
A verb, what it does, and the schema of its arguments and its answer. A name alone is not callable by
something that has never seen the mesh before, which is the whole population this surface exists for.
The protocol a seat carries today is three lists of bare verbs
([ADR 0129](../../02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md)), and it must widen to
carry the rest. Two constraints on that widening:
- **It lives in the mesh's records, not in the control plane's binary.** Today a seat's protocol comes
from compiled defaults, merged in as a row is read, because the seat rows never gained the columns.
Discovery that reads a binary is discovery that disagrees with the mesh the moment the two are on
different versions.
- **The schema is stated in the form an agent protocol already uses**, so nothing translates between a
seat's idea of an argument and the caller's. A translation layer would be a second definition of
what a tool is.
## 3. Holding a seat means serving its tools
A module may not occupy a seat unless it serves every verb that seat declares. This joins the
conditions of holding that already exist — providing what the seat delivers, being assigned at the
seat's scope — and is refused the same way: at registration and at handover, naming the verbs that are
missing rather than the fact that something is.
A module knows which seats it claims, so knowing which tools it must serve is not a discovery problem
for the module: the seat says, the module implements, and anything beyond that is its own.
## 4. Addressing a node-scoped seat
A seat's subject is flat today — `mesh.seat.<seat>.<kind>.<verb>` — which is correct for a seat the
mesh has one holder of and wrong for the six node-scoped seats, where one subject would reach every
machine's holder and the holders' queue group would hand the call to whichever answered first. A
node-scoped seat's tool therefore carries the node it is asked of. Nothing about a mesh-scoped seat
changes.
## 5. Discovery
**What a role answers is a read.** The seats and their protocols are records, so the list is a query
against the mesh's own store: no call to a module in the path, nothing that has to be running, and an
answer that stays true while a holder is restarting or being replaced.
**What a module answers comes from the module.** Its definitions live in its code, so it is asked, and
the answer is as available as the module is — which is the right coupling for a tool that only exists
while that module does.
A caller therefore gets one list assembled from two sources, and the difference is visible in it: a
role's tool names a seat, a module's names a module. An agent that wants to survive a holder being
replaced binds to the first.
## 6. What serves this to an agent
A module the mesh assigns to the machine where the agent runs, holding a credential the mesh minted,
with authority derived from what it may call — not a program started by hand with a credential printed
to a terminal. The adapter itself already exists and is thin by design; what changes is that it stops
being something a person carries and becomes something the mesh runs, on a node, like everything else.
An agent's authority can then be role-shaped: *the forge's tools*, rather than a list of
module-specific names that changes the day the forge is replaced.
## 7. Versioning
A seat's tools are an interface and change like one. Additive within a version. A change that would
break a caller takes the version token the subject already has room for, and the two versions run side
by side until nothing is bound to the old one.
## How it is checked
- **A holder missing a verb cannot take the seat.** One test per condition of holding, as the existing
conditions have, and the live refusal names the verbs.
- **A verb nobody declared is a subject nobody may use.** The bus grants are derived from the seat's
protocol already, and the golden composition of the user list is what keeps that honest: a holder is
granted exactly the seat's verbs, a user of the seat exactly the publish side.
- **Discovery needs no running module.** The test for a role's tools reads records and asserts the
answer equals what the seats declare — if it needed a module up, it would not be a read.
- **Two nodes holding one node-scoped seat derive two addresses.** Checked by the same test as the rest
of the subject table.
## What this does not settle
- Which verbs each seat should serve. That is a decision per seat, and the reason to do it slowly: a
seat's tools bind every future holder.
- Whether a module's own tool definitions should also be recorded when a build resolves its manifest.
There is an argument for it — the mesh could then answer for a module that is down — and an argument
against, which is that a recorded copy of a live definition is a copy that can be wrong.
## References
- [ADR 0132](../../02-DECISIONS/0132-a-seat-carries-the-tools-its-holder-must-serve.md) — the decision this designs
- [ADR 0129](../../02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md) — a seat carries the protocol of its role
- [ADR 0095](../../02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md) — a tool call passes one process where an audit belongs
- [`26-the-seats.md`](26-the-seats.md) — what a seat is, how it is held and handed over
- [`25-the-bus-on-nats.md`](25-the-bus-on-nats.md) §7 — a person's account, their inbox, and the adapter
@@ -155,3 +155,17 @@ design document here, and get it back. That check fails today by design.
**What stands until then** is the signpost, and the honest description of it: reachable, not
surfacing.
## Where this stands, 2026-09-29
*Added in a grooming pass.* The knowledge base this record is about is the **predecessor's**, and it
is no longer reachable from anything: the surface that answered `recall_search` speaks the transport
the mesh removed at the cut-over
([issue 147](../147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md)).
So the sentence in `README.md` that this record catches — *these documents are still indexed into
the knowledge base* — is now wrong twice over: nothing indexed them, and there is nothing to index
them into. The record stays open, and its answer is no longer "index this repository somewhere"; it
is whatever the mesh grows as its own knowledge surface, if it grows one. Until then the honest fix
is the README, which should stop claiming a property nothing provides.
@@ -1,8 +1,8 @@
---
status: open
status: resolved
opened: 2026-09-22
located-in: []
fixed-by:
located-in: [mesh-catalog modules/gitea, mesh-controller internal/catalogue/declaration.go]
fixed-by: mesh-controller 7352c84, merged in #46 — a module is told its port in a container's environment too, as a file already was
amended-design:
---
@@ -39,3 +39,9 @@ the assignment happens to differ.
- Should composition refuse an environment value that names a port the module does not fix, the
way it refuses other claims a module cannot make?
- Which other modules write their own address, with a port, into their environment?
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* The forge's address follows a moved port the same way every other reader does. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -1,8 +1,8 @@
---
status: open
status: resolved
opened: 2026-09-22
located-in: []
fixed-by:
located-in: [mesh-controller internal/catalogue/declaration.go, mesh-catalog (every routed module)]
fixed-by: mesh-controller bdf965d (a route names the endpoint it serves) with `portOfEndpoint` and `AtPublishedPort` — the contribution carries the endpoint's declared port and the machine-side redirection is applied to it like any other
amended-design:
---
@@ -45,3 +45,9 @@ precisely because the predecessor holds the usual one.
keeping the mapping out of rendered configuration?
- What should refuse a declaration whose contributed route names a port nothing on that node
listens on?
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A route names an endpoint rather than a port, and the redirection that turns a declared port into the published one is applied to contributions too. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -1,8 +1,8 @@
---
status: open
status: resolved
opened: 2026-09-24
located-in: [mesh-controller module.json, mesh-host internal/apply]
fixed-by:
fixed-by: ADR 0142 — the mesh's own components are delivered as binaries on the machine, so the controller is a process; the delivery itself is issue 142
amended-design:
---
@@ -84,3 +84,21 @@ restart and run-to-completion semantics — so this would not need host-side wor
The two do not collapse into one. The controller is not a code-carrying sidecar, and `network: host`
is what makes the asymmetry visible here and nowhere else.
## Answered
*2026-09-29, in a grooming pass.* This asked a question rather than reporting a defect, and the
question was taken: [ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md)
decides that the mesh's own components — the host, the controller, the catalogue, the builder, the
vault — are **binaries on the machine**, delivered by the mechanism
[ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) built, and that
third-party software (the store, the registry, the broker) stays a container because an image is the
right way to carry somebody else's build.
So the operating experience this record was written from — every mutating command reached through
`docker exec mesh-controller` — is answered, and answered against the container.
**The delivery is a separate matter and is not this record's.** Step 1 of it is built and no
component travels yet; that is
[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md).
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-26
located-in: [mesh-sdk src/provisioner, mesh-catalog modules/redis]
fixed-by:
fixed-by: mesh-catalog bbda88c, merged in #84 — every credential provider says whether it still holds a consumer
amended-design:
---
@@ -60,3 +60,9 @@ checks it after the first pass.
instance and leaves the gap for the others.
- Where does the record of what was applied live, if not in memory? ADR 0114, still
proposed, puts rotation state with the vault. The same place may answer this.
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A provisioner asks the backend what is there rather than trusting what it remembers doing. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -1,5 +1,6 @@
---
status: located
status: resolved
fixed-by: mesh-host 1cb8953 and fdc768c — the mesh writes into a marked block of a text file instead of over it
opened: 2026-09-26
located-in: [mesh-controller internal/catalogue/facts.go, mesh-host internal/apply]
---
@@ -71,3 +72,9 @@ private network loses that name too.
- The host's file resource supports `into: "json"` only; anything else is a whole write.
- `node show <node>` on the adopted workstation: `holds file /etc/hosts
mesh-wireguard.fact-node-names`, original kept.
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A shared hosts file keeps every line that is not the mesh's. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -1,7 +1,7 @@
---
status: open
status: located
opened: 2026-09-26
located-in: [mesh-controller, mesh-catalog step-ca]
located-in: [mesh-catalog ca-trust]
---
# 129 — nothing makes a machine trust the mesh's own certificate authority
@@ -0,0 +1,45 @@
# Diagnosis
*2026-09-29.*
## What was ruled out
**That something already carries the root and it is only misplaced.** It does not. The authority
serves its root at a path beside its ACME directory, and the one thing that fetches it — the route
proxy — puts it in a directory of its own and hands it to one program. Nothing has ever written
into a machine's trust store. Measured on three converged machines: the anchors present are the
predecessor's authority and a developer tool's local root, and on the machines where the
predecessor's was deliberately removed, every internal name fails verification.
**That the private network could carry it, the way it carries the registry's trust.** That is what
the report proposed, and it was rejected on consideration rather than on difficulty
([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md), option 1): being on
the network is what makes the registry *reachable* and is therefore the right trigger there, while
trusting an authority is a separate fact from being able to reach it. The anchor's directory and
the command that refreshes the extracted bundles are also one operating system's difference, which
is the host's half of the mesh and not the controller's.
**That it needs a new host resource type.** It does not, today. A file and a service say the whole
of it, which the packet filter already proves. The primitive becomes the right answer when a second
operating system is in play, and not before.
## Where it belongs
A module in the catalogue: it requires `internal-acme-ca`, fetches the root over the mesh's own
network, installs it as a trust anchor, refreshes the machine's bundles, and — because being
unassigned stops its unit, and stopping the unit is what undoes it — takes both away again.
The owner is therefore `mesh-catalog`, module `ca-trust`, and nothing in the control plane.
## The module exists, and this stays open until a machine holds it
*2026-09-29.* `ca-trust` is in the catalogue and merged
([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md)), and what it renders
is checked in the control plane's own suite: the script fetches from the authority it was bound to,
and the unit runs it both ways.
**No machine has been assigned it, and nothing has verified a name because of it.** The bed written
for that cannot run ([issue 146](../146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md)),
and the live mesh has not been given the module. So the symptom this record opened on — every
internal name failing verification on every machine — is still true everywhere, and the record stays
`located` until it is not. Closing it on a module that exists would be closing it on an intention.
@@ -1,5 +1,6 @@
---
status: located
status: resolved
fixed-by: mesh-host 3112c88 — undeclaring gives a unit back the state it was found in, and removes only a process the mesh made
opened: 2026-09-27
located-in: [mesh-host internal/apply/apply.go (remove)]
amended-design: 02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md
@@ -15,7 +16,7 @@ found that the host's `remove` path stops every `service` resource that is no lo
delete". `store.Orphans` matches by id alone. So any of these stops the unit:
- the module is unassigned — by mistake, or to switch it for another;
- the node is sent a deliberately-empty declaration ([issue 127](../127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md));
- the node is sent a deliberately-empty declaration ([issue 149](../149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md));
- a later catalogue version renames the resource's `id`.
That is right for a service the mesh brought into being. It is wrong for a unit the mesh
@@ -64,3 +65,9 @@ something to settle in passing.
The unassign preview is partly answered — the host's plan names each unit it will stop — and the
controller's side is left open.
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* Undeclaring no longer stops a unit the mesh only reloaded or only kept running. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -0,0 +1,78 @@
---
status: resolved
opened: 2026-09-28
located-in: [mesh-controller module.json]
fixed-by: mesh-controller — the control plane's module declares a run-once `migrate` step before its server, which is the shape ADR 0052 prescribes for exactly this. A step's record of having run is the digest of its declaration and the image is part of that digest, so a new build of the control plane re-runs it; and because a run-once step gates what the declaration places after it, a migration that fails stops the new server from starting at all rather than letting it run against a schema it does not have.
amended-design: 03-DESIGN/01-to-be/32-what-a-module-declares.md
---
# 133 — The control plane's schema is migrated at birth and never again
## What was observed
On 2026-09-28 at 08:17 the control plane was replaced, by the mesh's own upgrade path, with a build
whose code writes a column that a migration **in that same build** creates. Nothing ran the migration.
For the next three quarters of an hour the mesh built things and recorded none of them. Every build
answered:
> ERROR: column "built_contexts" of relation "build" does not exist (SQLSTATE 42703)
and that sentence went only to whoever happened to be waiting on a build's reply. The overview kept
saying the mesh was fine. The builds themselves worked — images were built and published — so the
registry filled up with artifacts the mesh has no record of, and the graph stopped learning without
anything saying so.
The schema was created once, at genesis, by an action in the foundation bundle that runs the same
binary's `migrate`. Nothing runs it again. The mesh has updated its own control plane many times since
that bundle, and every one of those updates carried whatever migrations the new build brought and
applied none of them. This is the first time a build needed one.
## Why it matters beyond this instance
**The schema and the code that needs it ship as one artifact and are applied by two mechanisms, only
one of which is automatic.** A module's version is atomic everywhere else in the mesh — the manifest,
the image and what the machine runs move together. Its schema did not, so "the mesh updates itself on
a push" was true of the code and false of what the code needs.
**The failure is quiet exactly where quiet is worst.** A build that cannot be recorded is a build that
happened and left no trace, which is the fault [issue 050](../050-the-catalogue-knows-nothing-built-before-it/00-report.md)
and [issue 131](../131-nothing-tells-the-mesh-a-source-moved/00-report.md) are both about. The mesh
has three mechanisms for noticing a module is behind its source and none for noticing that what it
recorded was refused.
**The shape was already decided, and the control plane was the one module that did not use it.**
[ADR 0052](../../02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md) says a run-once container is a
step the host runs to completion before whatever the declaration places after it, and names migrating
a schema as the case it exists for. The genesis code's own comment says a manifest may name its image
in more than one resource — "a migrate step beside the server". The control plane's manifest had no
such step; it went straight from a state directory to the server.
## What is still true
**Additive migrations are load-bearing, not a style preference.** The step runs before the *new*
server starts, which means the old binary briefly runs against the new schema. A migration that
removes or renames something would break the running control plane in the window between the two.
**A hand-written step is one the next module forgets**, which is why this fix is not where the matter
ends: [ADR 0135](../../02-DECISIONS/0135-a-module-version-prepares-its-state-before-it-runs.md) makes it
derived and puts it where an author works: a module version declares an entrypoint that prepares its
state, and the mesh composes the gated work from it, so the control plane stops being the only module
that had to remember. That record also settles the level question HAL answered with stages — a consumer
is a module on a machine, so the scope of preparation is the scope of the state — and
[ADR 0134](../../02-DECISIONS/0134-the-mesh-says-what-it-applied.md) answers the second open question
below: what a machine applied, and what it refused, become facts on the bus rather than a line in a log.
**The mesh now has two shapes for one problem.** The catalogue module migrates its own schema in its
own code when it starts; the control plane migrates in a step the host gates on. Both work and the
reasons differ — a module that owns its store entirely can do it at start, while a step is visible in
the declaration and refuses to let a broken upgrade serve. Which one the mesh should standardise on is
a decision, not a fix, and it is not made here.
## Open questions
- Should a module be refusable at registration when it ships migrations and declares no step and no
other way to apply them? The mesh can see both halves.
- Should a record the store refuses reach the overview? Today the only reader of that failure is
whoever asked for the thing that failed, and for an event arriving on the bus there is no such
person.
@@ -0,0 +1,66 @@
---
status: open
opened: 2026-09-28
located-in: [mesh-catalog, mesh-controller internal/catalogue]
fixed-by:
amended-design:
---
# 134 — A definition may still name the mesh, and the check that would say so does not exist
## What was observed
[ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) says a module
definition names no node, no mesh and no host path, and states how that is checked:
> A catalogue test finds no domain name in any definition value.
There is no such test. Run by hand on 2026-09-28, across the 72 manifests in the catalogue, the
question it asks has 15 answers. They are not all the same kind of thing, and the difference matters
more than the count:
**Values the mesh acts on** — seven:
| module | where | what it names |
|---|---|---|
| keycloak | `env.KC_HOSTNAME` | this installation's public name for itself |
| minio | `env.MINIO_BROWSER_REDIRECT_URL` | the same, for its console |
| invoicing | a resource's `image` | a named registry rather than the mesh's artifact store |
| builder | `build.artifacts[].context.repository` | the forge, by URL |
| route-proxy | `build.artifacts[].context.repository` | the forge, by URL |
| route-adapter | a resource's `content` | a proxy's dynamic configuration |
| novox.be | `module` | the module is named after the domain it serves |
**Prose** — eight, in `listens[].why`: de-spiegel, mailu, n8n, only-office, photos, photos-eef,
photos-filip, portainer. Each explains what a port is for and mentions the public name it is reached
by. Nothing reads these; a check written as a string search would report them, and reporting them as
violations of the same rule would be wrong.
## Why it matters beyond this instance
**An unenforced rule is indistinguishable from a wrong one, and costs more, because people believe
it.** The record says the mesh is name-agnostic, four design documents rest on that, and a reader
checking whether it holds finds that it does not — in the places that matter most. The two forge URLs
are what a build reaches into for its source; the two hostnames are what a service tells a browser
about itself.
**It is the difference between a mesh and this mesh.** A definition carrying `novox.be` is a
definition that can only be installed here. The whole point of the rule is that the same catalogue
raises a different mesh with a different name, and today seven modules would need editing to do it.
**And the shape of the fix is not the same for each.** A public name is an operator's choice about an
assignment, which ADR 0112 already provides for; a forge URL should be a path on the git seat
([ADR 0111](../../02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md)); an image from a named
registry is a question about the artifact store, not about naming. Counting them together would hide
that.
## Open questions
- Does a domain in a `why` string break the rule? It is documentation the mesh never reads, and a
check that cannot tell the two apart will either pass things it should catch or fail things nobody
should change.
- Where does a service's public name live, concretely — a setting on the assignment, or a fact the
mesh composes from the node's domain? ADR 0112 says a requirement the mesh resolves; the two
hostnames above are the first real cases.
- Should a build context name a repository on the git seat rather than by URL, and if so, what does
that mean for a context in *another* mesh's forge?
@@ -0,0 +1,70 @@
---
status: resolved
opened: 2026-09-28
located-in: [mesh-host internal/apply]
fixed-by: mesh-host — a container's mesh names are part of the spec digest the host compares, sorted so the digest does not move for a reordering. A container whose names moved is now recreated like a container whose image moved, and the test fails against the previous behaviour.
amended-design:
---
# 135 — A container's mesh names are not compared, so a moved address is never noticed
## What was observed
One container on this mesh had been restarting every thirty seconds for five days — 2286 times — and
the mesh reported the machine as doing what it was told.
Its logs said its database connected and then a query timed out. The database was reachable: the same
query from the same network, with the same credential, answered in milliseconds. What differed was the
name. Inside that container, `novox.internal` resolved to `10.42.0.1`; in every other container on the
machine it resolved to `10.10.0.1`. The mesh's overlay range had moved, and this container still held
the old one:
```
umami created 2026-09-23 novox.internal:10.42.0.1
mesh-catalog created today novox.internal:10.10.0.1
```
A container resolves other machines and public names through the entries the mesh gives it when it is
created, and nothing re-reads them afterwards. The host compares a container against what was declared
by a digest of its spec — image, name, environment, ports, volumes, arguments, resolver, address, and
what it reads — and **the mesh's names were not in it**. So this container matched what was declared,
was left alone, and kept an address that had not existed for five days.
Forty-eight other containers had current names. Not because anything corrected them: each had been
recreated for some other reason — a new image, a changed file — and picked up the current roster on the
way. This one's image is an upstream release that had not moved, and nothing else about it changed, so
nothing ever recreated it.
## Why it matters beyond this instance
**It is the exact fault [issue 045](../045-a-container-keeps-the-values-it-started-with/00-report.md)
named, in the one field that was left out.** That issue is why the digest carries what a container
reads: "a container whose configuration had since been rewritten compared equal and was left alone —
running values the machine no longer holds, while every check reported success." The same sentence
describes this, with *names* in place of *files*.
**The failure is invisible in exactly the way that matters.** The container runs, so the machine
reports it applied. It restarts, but a restarting container is a normal sight during an upgrade. The
only account of the fault is inside the container's own log, in the words of the application rather
than of the mesh — and what it says is that a query timed out, which points at the database.
**And it is most likely to bite what changes least.** Every container that is rebuilt often repairs
itself by accident. The victim is the module whose image is stable — which is to say, the module that
was working fine.
## What was done
The mesh's names are part of the digest, sorted so the digest does not move for a reordering nobody
made. A container whose names moved is now recreated exactly as one whose image moved.
The first apply after this recreates every container that carries mesh names — one restart each,
already the price the mesh pays for any image update — because their recorded digests predate the
field.
## What is still true
The mesh gives a container its names at creation and has no way to change them in place. That is the
container runtime's shape, not a choice; the answer is to recreate, which is what this does. A module
that would rather re-read a roster from a file can already ask for one as a fact
([ADR 0120](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md)) and restart on
it.
@@ -0,0 +1,92 @@
---
status: resolved
opened: 2026-09-28
located-in: [mesh-catalog modules/fail2ban]
fixed-by: mesh-catalog — the intrusion-prevention module bans through an action it ships itself, already in use on every machine, instead of naming a firewall front-end two of them do not have. The instance is closed; the class in "What is still true" is not.
amended-design:
---
# 136 — A module may name a program the machine does not have, and everything reports success
## What was observed
Two machines were given the intrusion-prevention module on 2026-09-28. Both refused to start it:
```
ERROR Failed during configuration: Have not found any log file for 'recidive' jail.
ERROR Async configuration of server failed
fail2ban.service: Main process exited, code=exited, status=255/EXCEPTION
```
The jail that bans whoever keeps coming back reads the service's *own* log, and the service checks
every jail's log file while it configures itself — before it has created that log. The module
declared the jail and shipped the rotation for that log, and never declared the log. On the two
machines where it had run for years the file was simply there, so nothing had ever noticed.
That failure was loud. Fixing it uncovered a second one in the same module that is not.
The module's defaults named `ufw` as the way to ban an address. Two of these four machines have no
`ufw` — they filter with nftables — and nothing checks that until an address is banned. Asked to ban
a documentation address on such a machine, the service accepted the instruction, counted it, ran the
command, and wrote this to a log nobody reads:
```
ERROR ... -- stderr: '/bin/sh: line 5: ufw: command not found'
ERROR ... -- returned 127
ERROR Failed to execute ban jail 'sshd' action 'ufw' ... Error banning 192.0.2.99
```
No rule existed afterwards. Throughout, the unit was `active`, the module was applied, and the
machine's report said so. **A machine had been added to the mesh's intrusion prevention, reported as
protected, and was banning nobody.**
## Why it matters beyond this instance
**The two faults are the same mistake with opposite symptoms.** Both are the module assuming
something about the machine — a file that happens to exist, a program that happens to be installed.
One stopped the service, which anybody notices. The other left it running and empty, which nobody
does. A mesh that only catches the loud one is a mesh whose coverage is unknown.
**"The unit is running" was taken for "the module is doing its job".** That is the only health a
service resource has. It is the right answer for most modules and it is silent for any module whose
work happens later, on an event — a ban, a renewal, a backup, a notification. The report cannot
distinguish "protecting this machine" from "installed and inert".
**And it is exactly the naming rule, one level down.**
[ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) says a
definition names no node, no mesh and no host path, because the same definition has to raise a
different mesh. `ufw` is not a node name, but it is the same class of assumption: a value the module
cannot know, true on some machines and false on others, written as though it were a constant. The
module already knew how to do better a few lines away — the mesh's own address range is named there
as something the machine fills in.
## What was done
The module declares the log its own jail reads, created once and never touched again, since what
grows in it is the service's and the rotation the module already ships is what keeps it small. And
it bans through the action it ships itself, which every machine here can run, which was already in
use by the other jail on all four, and which covers a container's published port as well as the
host's own.
All four machines now run it, with both jails, and a ban lands on each — verified by banning and
unbanning a documentation address on every one.
## What is still true
**Nothing would have caught either fault before it shipped.** The control plane reads a manifest, not
a machine; `ufw` and `/var/log/…` are strings in a file it has no way to evaluate. The host could in
principle be asked whether a declared program exists, but no resource says "this file names a command
that must be there", so there is nothing to check.
**Two machines' bans from before this are stale rules in the old front-end**, which the service no
longer knows about and will never lift. They reject two addresses for ever. Harmless, and a reminder
that changing how a module enforces something leaves what it already enforced behind.
## Open questions
- What does a service resource's health mean for a module whose work is event-driven? A unit being
active is the weakest claim available, and four of this mesh's modules are of that kind.
- Should a declaration be able to say that a resource depends on a program, so the machine can refuse
what it cannot carry out rather than reporting success?
- Where should the packet filter a module bans through come from — the module's own choice, as now,
or the seat that owns the machine's filtering?
@@ -0,0 +1,74 @@
---
status: resolved
opened: 2026-09-28
located-in: [mesh-controller internal/catalogue]
fixed-by: mesh-controller — a machine says which networks it routes and the derived filter forwards them, their guests keeping address and name service ([ADR 0137](../../02-DECISIONS/0137-a-machine-says-which-networks-it-routes.md)).
amended-design:
---
# 137 — Converging a machine cut off its own guests, and nothing said so
## What was observed
A workstation was flipped from adopted to converged, so the mesh's derived filter replaced what was
there. The flip reported success, the machine reported that it had applied its declaration, and every
surface of the mesh read green.
A container on one of that machine's networks could no longer reach anything:
```
192.168.64.2/20
OUTBOUND BLOCKED
```
Five of the machine's container networks were affected, and every network its test beds create. The
reason is in the filter's forward chain, which denies by default and then allows two ranges:
```
ip saddr 172.16.0.0/12 accept # the container runtime's bridge networks
ip saddr 192.168.128.0/17 accept # the networks its compose files are given
```
Those two are constants in the controller. The machine's guests were allocated from neither: its
compose networks from other parts of `192.168/16`, and each test bed a fresh `10.x/24`. So the rules
were correct for a machine whose runtime was left at its defaults, and a guess on this one.
Two further things were closed by the same flip, and for the same reason nobody saw them: a guest asks
its host for an address over DHCP and for names over DNS, both of which arrive at the input chain,
where no module had declared them.
## Why it matters beyond this instance
**The preview could not have warned.** It lists what the machine reported as *listening*, and says so
honestly: it ends with a line that traffic the machine routes is "not previewed". What it did not say
is that routing was about to be denied by default, or which ranges would survive. An operator reading
a 350-line preview approves what it shows.
**It is the second time today that a constant stood in for something the mesh cannot know.** The
intrusion-prevention module named a firewall front-end two machines do not have
([issue 136](../136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md)), and the
filter names the address ranges one runtime happens to use. Both were true where they were written and
silently false elsewhere.
**And the code already knew.** The comment above those two lines says a machine configured otherwise
"needs this to say so — which is a thing the mesh cannot derive and a reason this list is named here
rather than computed". The gap was documented at the point where it was introduced, and the way to say
it was never built. A comment naming a missing mechanism is a rule that is not enforced.
## What was done
A machine says which networks it routes; the filter forwards them and admits their guests' address and
name service. Added to the runtime's defaults rather than replacing them, so a machine that names one
range keeps the others. Node-level, because the machine routes them and the module that loads the
filter may be replaced. The converge preview now says what a machine routes, and what it will keep
forwarding if it says nothing.
## What is still true
**The flip is still the moment a machine's unmanaged services close.** That is what converging means
and the preview names each one. This issue is not about the ports that were meant to close; it is about
the ones nothing could name.
**Egress is still not previewed per network.** The preview says which ranges will be forwarded, not
which of the machine's guests sit inside them. Deriving that would need the machine to report its
bridges, and a bed's bridge does not exist until the bed runs.
@@ -0,0 +1,56 @@
---
status: open
opened: 2026-09-28
located-in: [mesh-controller internal/catalogue, mesh-catalog]
fixed-by:
amended-design:
---
# 138 — Two modules claim one seat and are not interchangeable, and nothing says so
## What was observed
Three modules claim the node-scoped uplink seat: one for each network manager a machine here might
run. [ADR 0117](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md) gives each of them the same
job — ask the manager the machine already runs to leave the resolver file alone and to leave the mesh's
interface alone — and deliberately keeps the machine's own links out of the mesh's hands.
A seat means one holder and an interchangeable holder. These are interchangeable in what they *ask*
and not in what they *do*:
- None installs, enables, starts or stops the manager. That is on purpose: stopping it takes every
link down, including the mesh's own way in.
- None carries an address, a route or a wireless credential, for the same reason.
- **Nothing checks that the module holding the seat names the manager the machine is actually
running.** Assigning the systemd-networkd holder to a machine running NetworkManager writes a file
for a daemon that is inactive and disabled, the seat reports held, and the two things the seat
exists to arrange are arranged for nobody. NetworkManager goes back to rewriting the resolver file
on every lease, which is the failure the module's own comment describes.
The machine reports which service manager and which units are active, so the fact needed to catch this
is already in the report the mesh holds.
## Why it matters beyond this instance
**A seat is the mesh's promise that a role is filled.** If the holder can be a module for software
that is not running, the seat says a role is filled while nothing fills it — and the surface that
would tell an operator says "held".
**It is the same shape as two faults found the same day.** A module named a firewall front-end the
machine does not have ([issue 136](../136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md)),
and the filter named address ranges one runtime happens to use
([issue 137](../137-converging-a-machine-cut-off-its-own-guests/00-report.md)). Each is a claim about
the machine that nothing on the machine checks.
**And it decides whether the seat is worth having.** Either the holder must match what the machine
runs, which is a condition the mesh can check from the report it already has, or the holders must be
able to switch the manager, which ADR 0117 refuses for a reason that has not changed.
## Open questions
- Should a seat's conditions of holding include a capability the machine reports, so a holder naming
absent or inactive software is refused rather than recorded?
- Is "the uplink" one seat at all, if its holders are three dialects of the same two requests? The
alternative is one module that speaks whichever dialect the machine needs, chosen from the report.
- What should happen on a machine that switches manager afterwards? The seat would then be held by the
wrong module, and the machine is the only place that knows.
@@ -0,0 +1,51 @@
---
status: open
opened: 2026-09-28
located-in: [mesh-controller internal/catalogue]
fixed-by:
amended-design:
---
# 139 — An internal route name resolves to the consumer's node, not the one that serves it
## What was observed
A module that requires a route is given two names: a public one composed under the serving node's
domain, and an internal one composed under the consumer's own machine — `<label>.<node>.internal`.
The two are published differently:
- The **public** name is written into every machine's hosts file at the address of the node whose
proxy answers it. The mesh computes that deliberately, so any container resolving a routed name
reaches the proxy.
- The **internal** name is resolved by the machine's own resolver, which answers every name under
`<node>.internal` with that node's address — the consumer's, because the name was composed from it.
Where the proxy runs beside the consumer these are the same machine, which is every case on this mesh
today, and both names work. Measured on 2026-09-28: the internal name of a service on the control node
answers with a certificate from the mesh's internal authority, and the public name with one from the
public authority.
Where the proxy is on another machine they disagree. The internal name sends the client to a machine
that runs no proxy and has nothing listening on the port, while the public name sends it to the one
that does.
## Why it matters beyond this instance
**It is latent exactly where the mesh is heading.** `route` is provided mesh-wide precisely so a
module can be routed by a proxy on another machine. The first module assigned that way gets an
internal name that does not work, and the public one that does — with no error anywhere, because both
names resolve.
**A per-machine name is what an operator will reach for.** `<service>.<machine>.internal` reads like a
promise that the service on that machine is reachable there, and the wildcard makes every such name
resolve whether or not anything answers.
## Open questions
- Should the internal name be composed under the serving node, like the public one, or should it stay
the consumer's and be published at the serving node's address like the public name is?
- Is a per-node route holder the real answer — a proxy on every machine that serves its own names —
and if so, is `route` still one mesh-wide provision or a node-scoped seat with a mesh-wide fallback?
- What certifies the name in either case? The certificate is obtained by whoever terminates TLS, and
that is the question above in another form.
@@ -0,0 +1,79 @@
---
status: located
opened: 2026-09-28
located-in:
- mesh-controller internal/catalogue/manifest.go
- mesh-controller internal/catalogue/filtering.go
- mesh-controller internal/catalogue/declaration.go
- mesh-controller examples/route-proxy
- mesh-catalog (every routed module manifest)
fixed-by:
amended-design: 03-DESIGN/01-to-be/08-connectivity.md
---
# 140 — An endpoint's reach is not declared, so three mechanisms each decide it separately
## What was observed
Preparing to converge the mesh's control-node — the last machine still running the firewall it
had before the mesh — the question came up for one module: the forge serves git over ssh, and that
port must stay reachable from outside the private network. Where is that said?
The manifest declares the port with a source of `mesh`, so the derived filter would close it to
everything but the private network. Looking for the place an assignment says otherwise, there are
two per-node settings keys: one that gives a module's declared port a machine port, and one that
overrides a declared port's source. The second has exactly one caller — the function that builds
the node's filter rules. Nothing else in the control plane reads it.
A module's routed endpoint is declared somewhere else entirely: a route contribution naming a label
and a port. It says nothing about reach. The proxy composes a **public** name and an **internal**
name for every route it is given, and obtains a certificate for each from a different authority.
Measured on that machine the same day: an identity provider's public name signed by the public
authority for 90 days, its internal name signed by the mesh's own intermediate for 24 hours and
renewed daily. Both names exist, and both certificates, because the proxy makes every name it can.
No assignment asked for either.
So the forge's ssh endpoint has a firewall source and nothing else — no name, no certificate, and no
way to say it should be public other than a key the filter alone reads. And the forge's web endpoint
has two names and two certificates that nobody requested.
## Why it matters beyond this instance
**Reach is stated twice, in two vocabularies, in two places that cannot disagree out loud.** A port
may be exposed to anywhere while the module contributes no public route; a public route may be served
for a module whose own listen is private. Nothing reconciles the pair or refuses it. Each mechanism
is separately defensible and the combination is unstated.
**The vocabulary belongs to the filter, not to reachability.** *Public, internal, or both* cannot be
expressed. A source of `anywhere` is one rule on one chain; it says nothing about which names should
exist or which authority should sign them. So "this endpoint must not be public" has no way to be
written, and is therefore enforced by nothing — while a public certificate for that very name is
obtained automatically.
**An endpoint is not a thing in the model.** A module has ports, and separately it has routes.
Nothing binds a port to a name to a certificate, which is why three mechanisms each decide reach on
their own and none of them is wrong. This is
[ADR 0045](../../02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)'s fault
one level up: that record closed "a declaration that reads as a restriction and restricts nothing"
for the packet filter. Here the declaration is absent altogether and the mechanisms guess.
**It blocks the certificate work.** The open question recorded for certificates — a name that must
not be public needs either DNS-01 or the internal authority only — cannot be answered while no
assignment states whether a name should be public. Neither can expiry reporting, revocation, or what
happens to a name when a machine leaves: all of them need to know which names were *meant*.
## Open questions
- Should an assignment name each of a module's endpoints, bind it to a node-level port, and state
whether it is reachable publicly, internally or both — with the filter, the proxy's names and the
certificate authority all derived from that one statement?
- What is an endpoint that is neither routed nor certified? Git over ssh is public reach with no name
and no certificate; the model has to hold that without inventing one.
- Are the two existing settings keys the same statement, half-built? If so, is this a new declaration
or the completion of theirs?
- Does an internal-only endpoint get a certificate at all, and from which authority — and does that
settle [issue 129](../129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md), where nothing
installs the mesh's own root?
- Does declaring reach per assignment also settle
[issue 139](../139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md), where an
internal name resolves to the consumer's machine instead of the one serving the endpoint?
@@ -0,0 +1,88 @@
---
status: located
opened: 2026-09-28
located-in:
- mesh-controller internal/catalogue/filtering.go
- mesh-host internal/apply
fixed-by:
amended-design: 03-DESIGN/01-to-be/08-connectivity.md
---
# 141 — The forward chain does not follow the modules, though the modules declare their networks
## What was observed
[ADR 0137](../../02-DECISIONS/0137-a-machine-says-which-networks-it-routes.md), decided the same
week, gave a machine a way to say which networks it routes for its guests, because the derived
filter's forward chain had until then allowed two ranges named as constants in the control plane's
own source — the container runtime's default bridge pool, and half of the pool its compose files
are given.
Checking the last machine still to be converged, the same fault was found to be live there, and the
declaration needed to work around it turned out to be wrong in kind.
That machine hosts twenty-one container networks. Nine fall inside the runtime's bridge pool and
are forwarded. Twelve sit in the other private range, and **six of those fall below the lower bound
of the constant**, so the flip would have cut their guests off exactly as it did on the workstation
that produced 0137.
Naming a range to cover the six was the obvious move, and is what 0137 provides for. But of those
six networks, **four are networks the mesh's own modules declare** — they appear as network
resources in the node's plan, created by the host because a module asked for them — and **two are
leftovers of the predecessor**, compose networks of services the mesh does not run. A range wide
enough to keep the four would have forwarded the two as well: a firewall widened by hand to protect
networks that should not exist.
The mesh already knows which of the twenty-one are its own. It made them.
## Why it matters beyond this instance
**The node's configuration is supposed to follow the modules assigned to it.** That is the mesh's
founding shape — the machine runs modules, and its files, its filter and its accounts are composed
from what runs there. The forward chain is the one derived thing that does not: it consults two
constants and, since 0137, a list a person types. A module added tomorrow brings a network the filter
will not forward; a module deprecated leaves a range in the list that outlives it.
**A typed range cannot distinguish the mesh's networks from what was left behind.** It is stated in
addresses, and addresses are what the runtime allocates, so the only honest declaration is one wide
enough to include whatever else the runtime has handed out. The derivation is narrower than anything
a person can safely write, because it names networks rather than ranges.
**0137 rejected deriving this, and was right about what it rejected.** It considered deriving the
list from *what the machine reports* and refused, on two grounds: a test bed creates its bridge
between one declaration and the next, and a filter that follows whatever appeared on the machine is a
firewall that widens itself. Deriving from the **declaration** is neither. The set is known before
the network exists, because a module declared it; and it cannot widen itself, because only a network
some module asked for is ever forwarded. What remains genuinely for a machine to say is guests no
module declares — a test bed's pool — which is a much smaller residue than the list as it stands.
**The gap is invisible in the one place that should show it.** The converge preview lists what
*listens*, and routing is not a listener. It says in one line what the machine routes, and a reader
has to know the runtime's allocations to tell whether that line is sufficient. On the machine
measured here it read as though nothing needed saying.
## What was decided
*2026-09-28, later the same day.* The answer is not a better list. The question in the first open
item below — should the chain be derived from the networks the modules declare — was answered *no*,
after a converged machine's rendered rules were read: the chain blocks everything passing through the
machine and then allows its own guests back by listing their addresses. Every route to a correct list
fails, because the mesh has no position on a container reaching outward in the first place. The filter
now constrains what arrives from **outside** the machine and says nothing about what did not, and a
machine says which of its links face outside — one reported fact instead of a list. See
[ADR 0140](../../02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md), which
supersedes both 0137 and the first attempt at answering this.
## Open questions
- Should the forward chain be derived from the network resources the node's modules declare, with the
host resolving each declared network to its address the way it already resolves a container by name?
The controller cannot render the address itself: a module's network resource carries a name, and the
runtime allocates the subnet at creation.
- What remains of `node networks` once that exists — only guests no module declares, such as a test
bed's pool? And should it then be named for that, rather than for all routing?
- The runtime's own default bridge, which containers attach to when no module network is named, is
not a module's network. Is it derived from the machine, declared by the module that owns the
runtime, or left as the one constant?
- Should the preview say which of a machine's networks are the mesh's and which are not, so a range
that exists to protect a leftover is visible as such?
@@ -0,0 +1,76 @@
---
status: located
opened: 2026-09-29
located-in:
- mesh-host internal/upgrade
- mesh-host cmd/mesh-host
- mesh-controller (no build source for the host; no resource delivers it)
fixed-by:
amended-design: 03-DESIGN/01-to-be/05-the-node-host.md
---
# 142 — The host is the one thing the mesh does not deliver
## What was observed
A change to the host was merged and could not reach any machine without a person copying a file.
Checked on the mesh of four machines, 2026-09-29:
- **The host is not a build target.** Asked what had been built for it, the control plane answered
`nothing has been built for mesh-host`. A merge on the forge builds every changed module and the
control plane itself, because the control plane is a module. The host is not one, and nothing
builds it.
- **No declaration delivers it.** No resource kind names an executable to place on a machine, and
nothing on a machine fetches one.
- **The half that recovers from a bad host exists and is unused.** `internal/upgrade` can report that
the executable this process started from has been replaced on disk, and records which version last
completed a reconcile so a shell script can roll back a host that will not start. The launcher reads
that record and rolls back. But `Replaced()` is called by nothing except its own tests — the
recovery is wired and the delivery was never built.
- **Every machine runs a byte-identical binary, stamped by hand.** All four carry the same size and
the same timestamp, from the last time somebody built it on a workstation and copied it out. No
package owns the file.
## Why it matters beyond this instance
**The component that implements updating is the one thing not updated.** The mesh's stated shape is
that a push produces the right builds and they reach the machines running them with nobody asking. It
is true of every module and of the control plane. It is false for the host, which is what applies all
of them.
**It is a bootstrap problem being answered by a person.** The host cannot be an ordinary module
because the host is what applies modules; a module that replaces the thing applying it has to survive
its own replacement. That is a real difficulty, and the work already done — noticing that the
executable changed, recording a known-good version, a launcher that rolls back — is the hard half of
solving it. What is missing is the easy half, and its absence makes the hard half dead code.
**A hand-copied binary has no record anywhere.** Nothing says which version a machine runs, so
nothing can say a machine is behind, and the mesh's own account of itself — every machine current with
its source — cannot include the host. Four machines agreeing today is luck, not a property.
**And it silently gates any change that starts in the host.** A change that needs the host to report
something new cannot be rolled out by merging it: the control plane must wait for a person, and until
then it either refuses what depends on the new report or renders something wrong. That cost is paid by
every future change of this shape, and it was paid today.
## Open questions
- How is the host delivered without being applied by itself? A candidate shape: the host is built like
anything else, published as an artifact, and the *running* host fetches and stages the next one, then
stands aside — which is what `Replaced()` was written for and what the launcher's rollback already
covers.
- **Should this ride the bus, rather than becoming a mechanism of its own?** Everything else that
reaches a machine already does: a declaration is sent over it, a report comes back over it, and a
build announces what it produced on it, which is how a module's new version reaches the machines
running it. A host build announcing itself the same way, consumed by the host already running,
would make this the existing mechanism pointed at one more artifact rather than a second way of
delivering things. It would also give the machine somewhere to say which host it is running, on the
report it already sends.
- What records which version of the host a machine runs, so "behind" is answerable? Nothing does now.
- Does the host's version belong in its report, beside the other facts a machine states about itself?
- Who decides when a machine takes a new host — the mesh, on a build, or an operator per machine as
with converging? The rollback path means a bad host costs a reconcile rather than a machine, which
argues for the former.
- Does the same gap apply to the launcher and the units beside the binary, which are also files no
declaration names?
@@ -0,0 +1,102 @@
---
status: located
opened: 2026-09-29
located-in:
- mesh-host internal/apply/opening.go (retireFirewall)
- mesh-host internal/apply/apply.go (the condition it is called under)
fixed-by:
amended-design:
---
# 143 — Converging a machine does not retire the firewall it found, and says it does
## What was observed
The control-node was converged on 2026-09-29, the first machine with a found firewall to be flipped —
the two converged before it had none.
The preview said, and the flip repeated:
```
the found firewall (ufw) is disabled, never flushed: its configuration stays on disk
...
sent: the host loads the mesh's filter and disables the firewall it found
```
The mesh then reported the node `converged`, 372 resources applied, nothing failed. Afterwards, on the
machine:
```
systemctl is-enabled ufw -> enabled
systemctl is-active ufw -> active
```
*Corrected 2026-09-29, an hour later, from reading the host rather than the declaration.* **The first
account of this was wrong.** It said the declaration carries no resource that would disable the found
firewall, and that the sentence was printed by the command with nothing implementing it. The
declaration indeed carries no such resource — but the mechanism was never meant to be one. It is a
step in the host's own apply, `retireFirewall`, and it exists, is careful, and is strict: it refuses to
retire anything until it has read back from the machine that the mesh's own table is loaded, it records
the forward policies first so a half-done retirement can be retried, and it verifies ufw reports
inactive afterwards.
What is established is narrower and stranger than "nothing implements it":
- ufw was **active and enabled two minutes after the flip**, and the flip had reported the node
converged with 372 resources applied and nothing failed.
- The machine's own record now reads `disabled_by_mesh: true` — but it was written by a reconcile
*after* an operator disabled ufw by hand, roughly fifty minutes later. A reconcile found ufw already
inactive, asked it to be inactive, read that back, and recorded that the mesh had done it.
- So the step did not take effect at the flip, and the machine's record now says it did.
The candidates are named rather than chosen, because the evidence does not separate them: the step is
called only when the apply had no failures, and a skipped step is silent; the mesh's table is loaded by
a service in the same apply, so whether it was loaded *at the moment the step asked* is an ordering
question; and the host's own detail lines do not reach the journal, so what it decided is not
recoverable after the fact.
## Why it matters beyond this instance
**It is a stated behaviour that does not happen, reported as success** — the fault this repository
exists to catch, and
[ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md) states it as
part of what the flip *is*: "loads the mesh's derived filter in place of its refusal-only table, and
retires the found firewall by disabling it, never by flushing".
**It could only be found on the first machine that had one.** The two machines converged before this
had no firewall to retire, so the step had never run, and nothing reported that it had not. That is
the same shape as [issue 136](../136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md):
a step that is silent when it does nothing.
**The machine is left doubly filtered, which is not what either firewall describes.** Every base chain
at a hook runs and a drop in any is final, so the machine now enforces the *intersection* of the mesh's
derived filter and a rule set left by the system being replaced. Nothing is broken by that today —
measured from outside, mail, the proxy and git-over-ssh answer and the databases and admin interfaces
are refused — but the machine's behaviour is described by neither of the two things claiming to
describe it, and the stale set includes a rule for a broker that no longer exists.
**And returning the node to adopted would be wrong in the other direction.** ADR 0100 says that
restores the found firewall by enabling it again; enabling something that was never disabled is
harmless, but the mesh's belief about which firewall is in force has been wrong in both modes.
## Open questions
- Which side owns retiring it — a resource in the declaration, so it is applied and reported like
everything else, or the flip as an act? A resource seems right: the flip is otherwise entirely
expressed as one, and an act that only the command performs cannot be re-checked on a later
reconcile.
- What should a reconcile do if the found firewall is enabled again by hand, or by a package update?
Convergence is a state, so presumably re-disable it and say so.
- Should the preview say what it *will* do rather than what it does, until a step exists that does it?
The wording was read as evidence twice in one session.
- Is there a check that a sentence the mesh prints corresponds to something that happened? This is the
second time in one session that a printed claim and the machine disagreed.
- **Why did the step not take effect?** It is called only when the apply had no failures, and being
skipped is silent. The mesh's table is loaded by a service in the same apply, so whether it was
loaded when the step asked is an ordering question — and ADR 0100 makes loading it first a
precondition rather than an expectation.
- **A step that records the mesh as having done what an operator did is worse than the omission.** The
record now says the mesh disabled ufw. Nothing distinguishes "we did this" from "we found it already
so". Should it?
- Why do the host's own detail lines not reach the journal? Everything it decided during the flip is
unrecoverable, which is why this account has candidates instead of a cause.
@@ -0,0 +1,84 @@
---
status: located
opened: 2026-09-29
located-in:
- mesh-host internal/apply/opening.go
- mesh-controller cmd/mesh-controller (the converge preview)
fixed-by:
amended-design:
---
# 144 — A predecessor's rules outlive the firewall the mesh found, and the mesh cannot see them
## What was observed
The mesh reports one thing about a machine's existing filtering: `firewall found: ufw`. On the
control-node, ufw was never what filtered the traffic that mattered.
Measured on 2026-09-29, before the machine was converged:
- ufw filters connections *to the machine*. It does not filter connections to a container's published
port, which arrive on the forwarded path where the container runtime accepts them before ufw's
forward chains are reached. Around thirty ports were published that way.
- Every one of the mesh's own forwarded openings, converged through ufw, had matched **zero packets** —
fifty rules in that chain, none ever matched, while the chain itself had passed 1.6 million
established packets. The restrictions read as applied and were inert.
- What actually kept those ports off the internet was a chain the predecessor installed in the
container runtime's own pre-accept hook, allowing the deliberately public ports and the private
ranges and dropping the rest on the outward link. Confirmed from outside: the proxy answered, the
container manager did not.
- That chain exists only in the running kernel. The persisted rule file is the distribution's empty
default, and nothing on disk recreates the chain.
After the flip, the mesh's own filter is loaded and does cover the forwarded path, so the machine no
longer depends on that chain. But **the chain is still there**, and it is now the only thing refusing
two ports the mesh believes are open: the bus and the registry, which
[ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md) requires be
reachable from anywhere so a machine can enrol and pull before it has a private-network address. The
mesh's rendered filter accepts both from anywhere. From outside, both are refused.
## Why it matters beyond this instance
**"The firewall found" is a kind, and filtering is not all in one place.** The host identifies one
front-end and reports it. A machine can carry rules from several sources — the front-end's own, the
container runtime's, an intrusion-prevention chain, and whatever a predecessor installed directly —
and the mesh's account of what filters the machine names exactly one of them.
**So adoption's central promise was half-true in both directions.** What the mesh converged through
the found firewall on the forwarded path did nothing at all, and what did the work was invisible to it.
A machine was reported as filtered by a mechanism that was not filtering.
**And convergence cannot retire what it cannot see.** Even once
[issue 143](../143-converging-does-not-retire-the-firewall-it-found/00-report.md) is fixed and the found
firewall is disabled, this chain remains, silently narrowing the machine below what the mesh's own
filter says. A rule the mesh did not write, cannot list, and will not remove — which today breaks the
enrolment path the design guarantees.
**The safe direction is not the same as the correct one.** Being more closed than intended broke nothing
visible, which is exactly why it went unnoticed for as long as the mesh has been on this machine.
## What it cost, measured later the same day
*2026-09-29.* The predecessor's chain was removed, and something it had been carrying went with it. It
admitted the private ranges wholesale, which is how a container on the machine reached a port declared
for the private network — the mesh's own filter admits the machines' overlay addresses, and a container
comes from a bridge. Every module that reached another by the machine's own name had been relying on the
predecessor's rule without anybody knowing.
That is [issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
and it ran for eleven hours while the mesh reported the machine healthy. The filter is fixed. What this
adds to the account here is that "the machine is more closed than the mesh believes" was not the
harmless direction after all — it was harmless for everything reached from outside, and an outage for
everything reached from within.
## Open questions
- Should the host report every place the machine filters from, rather than one kind — the front-end,
the runtime's hooks, and any chain it does not recognise, named so a person can look?
- What should the mesh do about rules it did not write and does not understand? Reporting them seems
right; removing them cannot be, and leaving them silent is what produced this.
- Does an opening converged through a found firewall need a check that it can actually take effect?
Fifty rules matching nothing would have been visible from the counters at any point.
- Is the bus and the registry being reachable from anywhere still what the mesh wants on a machine that
faces the internet? The design says yes, for enrolment. It deserves asking on its own rather than
being answered by a leftover.
@@ -0,0 +1,119 @@
---
status: located
opened: 2026-09-29
located-in:
- mesh-controller internal/catalogue/filtering.go (fixed for this instance)
- mesh-controller (what status reports, and what it does not ask)
fixed-by:
amended-design: 03-DESIGN/01-to-be/10-delivery.md
---
# 145 — A machine reads healthy while its modules cannot reach each other
## What was observed
Converging the control-node closed every path by which a module on that machine reached another module
by the machine's own name. It ran for **eleven hours**. Throughout, the mesh answered:
```
4 machine(s), all doing what they were told, all heard from,
running what the mesh would send them, and every module current with its source
```
What was actually happening, from one affected module's own log:
```
Doctrine\DBAL\Exception: Failed to connect to the database:
SQLSTATE[08006] connection to server at "novox.internal" (10.10.0.1), port 6852 failed: timeout expired
```
6,154 of them, beginning at the minute of the flip. The web application accepted TCP connections and
never answered an HTTP request; a client waited 35 seconds and gave up. Confirmed from a throwaway
container on the machine: neither the store nor the forge was reachable on the machine's own address.
The cause is [issue 144](../144-the-predecessors-rules-outlive-the-firewall-it-was-found-as/00-report.md)'s
sibling and is fixed: a port declared reachable from the private network admitted the machines' own
overlay addresses, and a container on the machine comes from a bridge address, matching none of them.
What this issue is about is the eleven hours.
**Nothing the mesh reports would have shown it.** Every check the mesh makes passed, because every
check the mesh makes is about the relationship between the mesh and a machine:
- the machine applied what it was sent, and said so;
- its declaration digest matches what the mesh would send;
- every module's source commit matches what the mesh holds;
- every container the declaration names is running.
None of those asks whether a module can reach what it requires. The mesh knows precisely who requires
what — it composes the grants — and never checks that the grant works.
**Nor would an operator's usual look.** The ports were probed from outside and behaved correctly; the
routed services answered; a container's egress to the internet worked. Those are the paths a person
checks after changing a firewall, and all three were fine. The broken path was module-to-module over
the machine's own name, which nothing routine exercises.
## Why it matters beyond this instance
**A mesh that composes a dependency and never tests it can only report on itself.** Every provision the
mesh grants is a claim that a consumer can reach a provider. The mesh asserts that claim, delivers
credentials for it, and has no mechanism that ever finds out. "Every module current with its source"
is a statement about bytes, not about whether anything works.
**The failure was silent in the direction that hides longest.** A service that will not start is
noticed. A service that starts, accepts connections and then cannot reach its database serves errors
under a healthy-looking process, and the machine's own report says the container is running — which it
is.
**It is the same shape as [issue 136](../136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md),
one level up.** There, a module named a program the machine lacked and everything reported success.
Here, the mesh granted a provision the filter refused and everything reported success. Both are the
distance between a declaration and the machine, and in both cases the report was about the declaration.
**And the eleven hours are the measurement, not the bug.** The filter fault was one line and is fixed.
What is not fixed is that nothing in the mesh would have told anybody.
## What was decided
*2026-09-29, the same day, in two steps and the first was wrong.*
The first answer was [ADR 0143](../../02-DECISIONS/0143-a-consumer-verifies-the-grant-it-is-given.md):
the consumer verifies each grant from its own network position, because whether a caller sat in a
container changed whether it could reach the provider. **That difference was the fault**, and the record
is superseded. A verification mechanism would have reported this sooner and would not have prevented it,
and the part of it that was difficult — deciding which network position to check from — existed only
while the rule was wrong.
The remedy is [ADR 0144](../../02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md):
anything on a machine may call anything on it, said once rather than per service, and asked by the link
traffic arrives on rather than the address it carries. Everything should be able to call what runs on the
same machine, another machine's service exposed to the private network, and another machine's service
exposed publicly. The filter had the second and third and expressed the first as a list of addresses that
no container could match.
**And then the question this issue is actually about was answered on its own terms.**
[ADR 0145](../../02-DECISIONS/0145-a-module-checks-what-the-mesh-claims-is-reachable.md): a module on
every machine serves an endpoint of its own and dials every other machine's, from the position the
callers are in. Its probe is its own endpoint declared reachable over the private network, so it is
admitted by exactly the rule that governs every internally-exposed service and fails when that rule is
wrong — where a probe on a service every machine has would have passed for all eleven hours, because the
services every machine has are the ones never closed.
Adopted on its merits rather than as the remedy for a configuration error, which is what 0143 was and
why it went. The module is written and merged; it is not yet assigned, so every machine currently reads
unchecked.
## Open questions
- Should a grant be checked? The mesh knows the consumer, the provider, the address and the port, so a
reachability check is expressible — but from where: the consumer's machine, as part of a reconcile,
or the provider's?
- What would it cost to be wrong in the other direction? A check that reports a provision broken while
it works is worse than none, because it trains a reader to ignore the report. A provider restarting is
ordinary; a consumer between containers is ordinary.
- What should `status` say about a machine whose modules cannot reach each other? It currently has one
vocabulary for "heard from and current", and that sentence was true the whole time.
- Is there a cheaper signal than a probe? The affected module was logging the failure 6,154 times. The
mesh reads no module's logs and arguably should not — but something a module could *say* about its
own provisions would have surfaced this in minutes.
- Does the same blindness apply to the other direction — a provider that lost a consumer's grant and
is refusing it? Nothing checks that either.
@@ -0,0 +1,73 @@
---
status: located
opened: 2026-09-29
located-in: [mesh-host examples + internal/link, mesh-controller internal/broker]
---
# 146 — the foundation cannot be raised on the bus the mesh runs on
## What was observed
Raising a first node in the lab, to check a module against a real mesh, fails before any module is
reached. Two separate faults, in the two bundles that exist:
**The older bundle raises a control plane that cannot start.** It brings up the previous broker,
and the control plane it then starts says, once every few seconds, for ever:
```
mesh-controller: this control plane has no MESH_BUS_NATS, so it cannot reach the mesh's bus
```
That is the control plane being right. The mesh moved to one bus
([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)) and the
bundle did not. Every bed that raises a foundation raises this one, so every bed is in this state.
**The newer bundle, written for the new bus, stops one step earlier.** Its certificate step asks a
container to make the broker's certificate:
```
docker run --rm --entrypoint sh -v <the broker's tls volume>:/tls <the bus image> \
-c "test -f /tls/tls.crt || (openssl req -x509 ... )"
...
failed bus-certificate: running the action: docker exited 127
```
127 is *command not found*. The bus's image has a shell and no `openssl`; the previous broker's
image had both, which is why the step worked when it was written against that one. Substituting the
store's image — the only other image the bundle carries — does not help: it has no `openssl`
either. So the step as written cannot succeed with anything the bundle names, and the fault is not
one image's: **the bundle asks for a certificate to be made by a tool it never says must be there.**
Measured 2026-09-29 on a fresh lab machine, both bundles, from bare.
## Why this is here and not a note in the knowledge base
The mesh's own foundation is the one thing it cannot raise. Nothing reports that: the bundles are
files in a repository, nothing applies them but a person raising a node, and the last thing that
did was the hand-driven cut-over
([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md), whose
work was done on the machines rather than from a bundle). So the state where the mesh cannot make
another one of itself is reachable, and was reached, without anything saying so.
It is also load-bearing for everything else: a lab bed proves a claim by raising a mesh, so while
this holds, **no bed can run**, and every "checked in the lab" written from now on is a promise
against a suite nobody can execute.
## What would have prevented it
- **Something raising the foundation on a schedule, from the bundle, as it is written** — the
bundle is the mesh's own installer and nothing installs from it. A bed that raises a first node
is exactly that check, and it is the bed that cannot run.
- **A step naming what it needs.** The certificate step names an image and assumes a program inside
it. An action that said which tool it requires would have failed at the declaration rather than
at 127 on a machine.
## Evidence to carry into diagnosis
- `mesh-host examples/foundation-first-node.lock` — the previous broker, no `MESH_BUS_NATS`.
- `mesh-host examples/foundation-first-node-nats.lock` — the new bus; `bus-certificate` and its
`verify` both run `openssl` in the bus's image.
- The bus image the bundle pins has `sh` and no `openssl`; the store's image likewise.
- The lab rewrites a bundle's registry-prefixed third-party references to upstream ones for a
machine with an uplink (`test/integration/harness.ts`); the new bundle's bus reference needed
that rule added, which is done and is not this issue.
@@ -0,0 +1,184 @@
# Diagnosis
*2026-09-29, by raising a first node in the lab over and over and writing down each thing it hit.*
Not one fault. **Four, stacked**, each hidden behind the one before it, and every one of them the
same shape: a step that was right while the mesh ran on the previous broker and was never asked a
question again after the bus changed. Nothing had raised a foundation since, so nothing said so.
## 1 — the bundle's bus image is named for a registry that is gone *(fixed)*
The newer bundle pins `<a lab registry>/nats@…`, which resolves nowhere outside the lab that
raised that registry. The lab already rewrites the store's and the previous broker's references to
upstream ones for a machine with an uplink; the bus had no such rule because no bed had ever tried
to raise this bundle. Added (`mesh-lab test/integration/harness.ts`). The digest is the bundle's
own — what the registry served was a copy, so the same digest resolves upstream, and this is a
prefix being removed rather than a reference being replaced.
## 2 — the bus's certificate was made by a tool the bus does not have *(fixed)*
```
failed bus-certificate … docker exited 127
```
The step ran `openssl` inside the broker's image. The previous broker's image carried it; the bus's
does not — it is Alpine with a shell and no `openssl` — and neither does any other image the bundle
names, so there was nothing to substitute. **The program that needs the certificate now makes it**:
`mesh-controller broker certificate --into <dir>`, with `--check` as the step's verify. The
controller is already on the machine at that point (the schema step ran it) and needs nothing from
the image it writes into. Self-signed, as before and on purpose — a host pins this server's exact
certificate ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)) and at that moment
there is no authority to ask. Idempotent, because a second certificate is one every host that
pinned the first no longer believes. It runs `--user 0:0`: the volume is root's, and the control
plane's image runs as nobody, which is right for the long-lived server and wrong for a one-shot
writing into a fresh volume.
## 3 — enrolment dialled TLS at a bus that speaks first *(fixed)*
```
mesh-host: cannot reach the broker at …:5671: tls: first record does not look like a TLS handshake
```
Enrolment opened a raw TLS connection to check the pinned certificate before saying anything. NATS
speaks its own protocol and upgrades afterwards, so the handshake met a plaintext greeting. The pin
was never the problem: the same pinned configuration is handed to the client that presents the
token, and the verification runs inside *that* handshake — so what
[ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md) requires still holds, and holds
better, because the one-time secret is sent only after the certificate has been checked. The raw
dial is gone from the enrolment path and kept only as what its tests always proved: that a wrong
certificate is refused before a byte of application data is sent.
**Then, immediately behind it:**
```
mesh-host: this token is for the "" bus, and the mesh's bus is nats
```
The enrolment left the transport empty and meant *whatever the mesh runs today*, which was true
while two buses existed and became a refusal the moment one did. The host knows which bus the mesh
runs; it says so now.
## 4 — a first node cannot be let onto its own bus *(open, and this is the real one)*
```
mesh-host: cannot reach the bus at …:5671 as anchor: nats: Authorization Violation
```
The bus's user list is a file beside its configuration. The installer carries the first one — the
controller's own account at a bootstrap password — and **the controller composes every user after
that** (design 25 §6; the controller's own test asserts the carried list matches what it would
derive). On the running mesh that composition reaches the bus because the bus is a *module*, with
the list delivered to it the way anything is delivered to a module.
At genesis there is no module. The foundation's bus is raised by the installer, the control plane is
given no way to write beside its configuration — it mounts the certificate and nothing else — and so
the account a joining node needs cannot come into existence. **The first node cannot join the mesh
it just raised.**
That is not a line to fix in a bundle. It is the open half of the mesh delivering its own components
([ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md)) and of the
bus becoming a module: either the installer's bus is raised as the module the mesh will go on
managing, or genesis carries a user list that includes the first node's enrolment and the controller
takes over from there. Both are decisions, not patches, and both belong to the genesis step that was
deliberately left until last.
## 5 — the composed user list has to be placed by hand at genesis *(fixed)*
The account a token is the password of is **not recorded at all**: the composer names an enrolment
user for every machine with a live token, nothing minted a credential for it, and the composition
left it out as a user with no password. The comment above the issuing code already claimed
otherwise — *"the account is created before the token is handed over"* — which is how it went
unnoticed. Issuing a token now records that account, with the token's own secret as its password,
because that is the string the machine will present.
Placing it is the other half. The list reaches the machine running the bus in that machine's
declaration, which a machine that has not enrolled does not get, so at genesis it cannot arrive
that way. **The control plane composes and says what it composed** — `broker accounts`, to standard
output — and whoever is raising the machine writes it beside the bus's configuration and makes the
server re-read it. Twice, because two accounts come into existence at different moments: the
enrolment when the token is issued, and the machine's own when it enrols. A control plane that
wrote the file itself would have to know where the bus keeps its configuration and how to make it
reload, which is the module's knowledge and is what the module takes over on the first push.
With that, **a first node enrols against the bus it just raised** — measured, from bare, in the
lab.
## 6 — and is enrolled twice, keeping a credential the mesh has replaced *(fixed)*
```
mesh-controller: enrolled anchor
mesh-controller: enrolled anchor (the same second)
```
One `enrol` on the machine, two enrolments in the control plane. Each mints the node a fresh bus
password and returns it; the machine keeps the answer to the first, and the mesh keeps the hash of
the second. The machine then reconnects for ever as a user whose password the mesh rotated out from
under it — *authentication error - User "node.anchor"* on the bus, `Authorization Violation` in the
host's log, and a node that never reports.
What is ruled out: the host asking twice — it asks again only when the mesh says *try again*, and
a refused attempt is not logged as an enrolment. Redelivery by the consumer — there is one
consumer, its acknowledgement window is thirty seconds, and the handler is quick.
What is left: the client re-publishing when an acknowledgement is slow, which is what its defaults
do. That was addressed by giving the publish a message id derived from its own bytes, so the stream
discards the copy — **and the duplicate survived it**, so either the id is not reaching the stream
or the second copy is not a copy. This is where the trail stops.
Worth saying plainly: **the mint is the fragile part, not the delivery.** An enrolment answered
twice is survivable if the answer is the same both times, and it cannot be — the mesh keeps only
the hash, so a second answer is necessarily a different credential.
**Found, and it is not about enrolment at all** *(fixed)*. The bus's own counters settled it: one
message published, one held in the stream, one delivery, nothing redelivered — and the controller
enrolled the machine twice. So the handler ran twice on one delivery.
A push consumer delivers onto an ordinary subject, and **everything subscribed to that subject gets
a copy**. The controller holds a consumer called `controller` on CONTROL and another called
`controller` on EVENTS, and the delivery subject was derived from the consumer's name alone — so
both were `_DELIVER.controller`, the one process held both subscriptions, and every message from
either stream was acted on twice.
Enrolment is where it drew blood, because enrolling twice mints twice and the second credential
replaces the first. But it applied to **every report and every event the controller follows**, and
it is the kind of fault that leaves no trace: nothing is redelivered, no counter is wrong, the work
simply happens twice. The comment in the receiving code about a merge that ran the whole catalogue
five times over on 2026-09-28 is the same shape seen from the other end.
The stream is in the delivery subject now, because the pair is what identifies a consumer — the
server scopes a durable's name to its stream, and this subject was the one place that scoping was
dropped. A subscriber's permission gains the same shape, keeping the bare name so an existing
consumer keeps working until the controller's next assertion moves it.
*How it is checked:* the consumers the mesh asks for are asserted to deliver onto distinct subjects,
in the controller's own suite. Against a server it would be invisible, which is the point.
## Where it belongs
`mesh-host` (the bundle and the enrolment path) and `mesh-controller` (the certificate command, and
the composition that cannot reach the bus at genesis). Three of the four are fixed on branches; the
fourth is the genesis work.
## What made it slow, and what was changed so it is not
Six faults behind one another, each found by raising a machine and reading what it said. What cost
the most was not the faults:
- **Every bed's own instructions named the bundle that cannot work**, so the first three attempts
ended in a control plane crash-looping on a missing bus. They name the working one now.
- **A host binary built without its system** refuses everything it is given with *this host was
built for ""*, which reads like a broken bundle. The lab's README says so.
- **`make image` in the control plane had been broken for as long as its base was pinned**: the
Dockerfile's fallback is a Go older than the module asks for, and the pipeline never saw it
because the pipeline passes the declared base in. It reads the base from the manifest now.
- **Leaving the machine standing is what answers the question.** Every finding above came from
shelling in afterwards — the host's log, the bus's log, the file the bus was actually handed —
and none from the test's own output, which says only that nothing converged. The bed takes
`MESH_LAB_KEEP`, and the README says to reach for it first.
## What it cost, for the next person
Every lab bed still names `foundation-first-node.lock` in its own instructions, and that bundle
raises the previous broker with a control plane that refuses to start without `MESH_BUS_NATS`. Until
the fourth fault is answered and the two bundles become one, a bed runs with `MESH_LAB_BUNDLE`
pointing at the NATS bundle by hand, and stops at the enrolment.
@@ -0,0 +1,75 @@
---
status: located
opened: 2026-09-29
located-in: [nothing in the mesh — the surface in use is the predecessor's, installed on the workstation]
---
# 147 — the operator's tools still dial the bus that was removed
## What was observed
Every tool call an operator makes against the mesh fails, on every machine, with the same answer:
```
AMQP not connected — cannot reach hal/mesh@novox
AMQP not connected — cannot reach hal/mesh@shanks
```
The mesh moved to one bus and the previous transport was deleted
([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md), cut over
2026-09-28). The surface an operator drives the mesh through — the tool bridge their assistant
speaks to — still opens an AMQP connection, so it cannot reach anything. Not one node answers,
including the machine the operator is sitting at.
**What that leaves.** The mesh itself is healthy: nodes are current, modules run, the bus carries
the mesh's own traffic. What is gone is the way a person asks it anything. Every question — what a
node runs, what is assigned, what a module's settings are — has to be asked by opening a shell on
the machine and running the control plane's binary inside its container, which is precisely the
path the tool surface exists to remove, and which nothing checks, records or permits.
It also silently changes how work gets done: an assistant told to use the mesh's tools finds them
dead, falls back to `ssh` and `docker exec`, and the operator discovers the fallback rather than
the fault.
## Why this is here and not a note in the knowledge base
The rule is that everything happens over the mesh's bus. A surface that cannot reach that bus is
that rule enforced by nothing — and the mesh reports itself healthy throughout, because what broke
is not a module, a node or a provision but the thing standing outside asking them questions.
The same cut-over removed the assistant's long-term memory (`recall`, `memorize`) for the same
reason, which is why lessons from the last two days were written into this repository by hand.
## What would have prevented it
- **The tool surface as a consumer of the bus like any other**, so moving the bus moves it — rather
than a separate bridge with its own connection settings that nothing resolves.
- **A check that a tool call reaches a node**, run where the mesh's other checks run. Every check
the mesh makes today is about the relationship between the mesh and a machine; none asks whether
a person can ask it anything ([issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
the same shape one level out).
## Diagnosed at once, because the answer was in the configuration
**The surface is not the mesh's.** The tool server the operator's assistant speaks to is the
predecessor's brain, installed on the workstation and started as a local process, with the
predecessor's broker URL — `amqp://…@<the control node>` — written into the assistant's own
configuration. Nothing about it is a module: it has no manifest, no assignment, no seat, no account
on the mesh's bus, and the mesh has never known it exists.
So nothing regressed. The mesh removed a transport that this program still dials, and the program
was never part of the mesh to be moved. **The mesh has a tool model** — `tools` in a manifest, the
calls a seat accepts, the subjects a module answers on, and a control-plane command that asks one —
and **nothing publishes an operator-facing surface onto it.** What an operator uses is the thing
that came before, kept alive by a URL in a file.
That is the issue, and it is larger than a broken connection: the way a person drives this mesh is
outside the mesh.
## Evidence to carry into diagnosis
- The failure text names `hal/mesh@<node>`, so the bridge is resolving a node and then dialling the
old transport.
- It fails identically for the local machine, which rules out reachability and points at the
transport alone.
- The mesh's own traffic over the new bus is unaffected: nodes report, declarations apply.
@@ -0,0 +1,37 @@
---
status: open
opened: 2026-09-29
located-in: [mesh-controller cmd/mesh-controller]
---
# 148 — a manifest outside this catalogue has no check
## What was observed
A module's manifest is validated by a **test** — `internal/catalogue`'s suite parses every manifest
in the catalogue checkout beside it and fails on one it cannot resolve. That works, and it is how
several real faults were caught before a machine saw them.
It is available to exactly one repository: this one. Somebody describing their own application in
their own repository — the case
[ADR 0037](../../02-DECISIONS/0037-where-a-module-lives.md) calls *the case that matters most* —
has no check at all. They write a manifest, register it with a running mesh, and find out whether
it is valid when the mesh refuses it, or later, when a machine applies something that resolved and
should not have.
The same record asks for the answer: **a `module check` command on the control plane's binary**, so
a manifest is checked by the tool rather than by a test that imports the tool's internals.
## What would have prevented it
Nothing prevents this; it was noticed and left. ADR 0037 named it on 2026-09-01 and the record sat
`proposed` until 2026-09-29, so the missing half was never anybody's task.
## Evidence to carry into diagnosis
- `mesh-controller/internal/catalogue` — `ParseManifest` and `CatalogueProblems` are the check, and
both are internal.
- The catalogue-wide test is `TestEveryCatalogueManifestDeclaresWhatItMounts` and its siblings; they
take a path from `MESH_CATALOG`, so the mechanism is already path-driven and not repository-bound.
- `mesh-controller module add` refuses a bad manifest at registration, which is the same check far
too late: by then it is in a running mesh's records.
@@ -1,10 +1,15 @@
---
status: located
status: resolved
opened: 2026-09-27
located-in: [mesh-controller cmd/mesh-controller/push.go]
located-in: [mesh-controller cmd/mesh-controller/push.go, mesh-controller cmd/mesh-controller/sendable.go]
fixed-by: mesh-controller sendable.go and push.go — an empty declaration is sent carrying `owns_nothing`, and the host refuses an empty body that does not carry it
---
# A declaration that shrinks to empty is skipped, so the node keeps what it should drop
# 149 — a declaration that shrinks to empty is skipped, so the node keeps what it should drop
*Opened as 127 and renumbered on 2026-09-29: two records were given that number on the same day.*
*The other kept it, because three documents and three source files cite it by number and nothing
cited this one but a decision and a sibling issue, both corrected with this move.*
## What was observed
@@ -37,3 +42,17 @@ mean "own nothing", which the host already applies correctly when it receives on
On ace, one command drops it permanently (the corrected controller never re-composes it):
`sudo ufw delete allow 5671`. At ace's converge it would clear on its own.
## Closed
*2026-09-29, in a grooming pass.* Both halves are on `main` and both name this issue.
- The control plane **sends** it: a declaration that composes to no resources goes out with
`owns_nothing`, and `push` says *sent, not skipped*.
- The host **refuses an empty body that does not carry it**, so a truncated or mis-composed
declaration can never be read as "own nothing" — which is the failure the fix had to avoid while
making the empty case expressible.
Closed by reading the code rather than by watching a machine let go of a stray resource; the record
says so rather than implying a run.