Compare commits

..
Author SHA1 Message Date
jschoubben 04c9500b5b Group 2 is resolved: containers resolve, nothing is copied, a route's name says where it arrives
Issue 110's cause was not the filter: the runtime had never been told,
and the resolver dropped a query arriving on a bridge. ADR 0148 step 3
landed once it did (109, 151 resolved). ADR 0151 composes a route's
internal name under the serving node and drops the suffixed alias
(139, 157 resolved). Design 08 amended; a fact in 0148 corrected.
2026-09-30 14:56:43 +02:00
jschoubben 846c1f85f2 Merge pull request 'Issue 107 is resolved: a declaration carries its order' (#217) from issue/107-resolved into main 2026-09-30 12:13:57 +00:00
jschoubben 9eef0bd525 Issue 107 is resolved: a declaration carries its order
Hosts first, then the controller — a build and a push each, now that the
mesh delivers the host. The host refuses a lower sequence than it kept
and drains a batch by sequence rather than arrival; the controller
numbers each send under the node's hold, inside the signed bytes.

Measured: two pushes, sequence 2 in the kept declaration, counters in
the store agree, no machine reads as behind. That last one is the
subtlety: the mesh compares the digest of what it would send against
what it did, and a number changes the bytes, so the read-only comparison
composes with the last number sent rather than a fresh one.
2026-09-30 14:13:50 +02:00
jschoubben 6e08cdf3d6 Merge pull request 'Every machine self-updates, verified, and 107's gate has opened' (#216) from issue/142-self-update-on-every-machine into main 2026-09-30 11:51:53 +00:00
jschoubben 02f291a129 Every machine self-updates, verified, and 107's gate has opened
All four machines run a host the mesh built, published and delivered, the
last delivery unattended: each stood aside once for a genuinely newer
version and the delivered launcher started it. A following push that
delivered nothing new was applied and reported by every machine and stood
nobody aside.

The crossover needs one restart of the unit per machine, once, because
the running launcher executes from its own inode. Measured timing: three
seconds on the machine, 17-20 as the operator sees it, the difference
being the control plane composing before it sends.

107 is unblocked: a declaration field is now a build and a push.
2026-09-30 13:51:46 +02:00
jschoubben 9b14430d3f Merge pull request 'Issue 163: a delivered host stood aside on every push and reported nothing' (#214) from issue/163-a-delivered-host-stands-aside-on-every-push into main 2026-09-30 11:47:35 +00:00
jschoubben a4384f13d3 Issue 163: a delivered host stood aside on every push and reported nothing
Asked whether a newer host was delivered using the link-time stamp, which
every delivered host carries as 'development build' now that the version
comes from where the binary sits. Never matched, so it stood aside on
every push for ever; standing aside cancels the report, so the mesh never
heard from it. Read as healthy throughout.

The three-minute push wait made it invisible: a wait long enough to
absorb a whole apply is long enough to hide that the machine never
answered.
2026-09-30 13:47:28 +02:00
jschoubben a841e2c173 Merge pull request 'The host self-updates, and an archive cannot be undeclared' (#212) from issue/161-resolved-and-162-an-archive-cannot-be-removed into main 2026-09-30 11:16:44 +00:00
jschoubben 3c535ead31 The host self-updates, and an archive cannot be undeclared
161 resolved and verified on a machine: the workstation runs a host the
mesh compiled, published, delivered and started, applying declarations
and reporting the version it was delivered as.

The system it was built for comes from the artifact — the one thing a
toolchain takes from a module, which 0142 already allowed because the
target is a property of the artifact. The version comes from where the
binary sits, which 0142 decided and nothing had implemented.

Two mistakes on the way, both caught by reading the output rather than
the line that claimed success. A second -ldflags does not merge with the
first: the binary gained its system and lost -s -w, 12.2MB against 8.5MB.
And the delivered binary was named after its package, so the first
delivery was correct, reported success and was invisible to the launcher.

A delivered host that cannot apply is a machine the mesh cannot repair,
because the declaration that would fix it is the one it cannot apply. The
launcher's fallback is what made that an inconvenience instead of an
expedition.

162 is new and not about the host: an archive has no removal, so a module
using one can never be unassigned, and the attempt takes the whole apply
with it — the machine applies nothing else either. It is how undoing the
first delivery froze the workstation.
2026-09-30 13:16:37 +02:00
jschoubben 4cf941d858 Merge pull request 'Self-update works, and a delivered host is one fact short of usable' (#211) from issue/161-a-delivered-host-has-no-link-time-facts into main 2026-09-30 10:28:08 +00:00
jschoubben 5042ffd8d3 Self-update works, and a delivered host is one fact short of usable
The loop closed on the workstation: the version landed, the launcher was
replaced, the running host stood aside, and after one restart the launcher
started a binary the mesh had compiled, published and delivered.

The launcher goes as a file resource rather than inside the archive, and
that is the safety rather than a preference. A file is written atomically,
so the running launcher keeps the inode it started from; an archive writes
in place with truncate and would cut a script a shell is reading. The
manifest carries a second copy and a test refuses any drift from the one
in packaging.

Then it would have refused the first declaration it was asked to apply.
The Makefile links in two facts the mesh's toolchain does not, on purpose,
and one of them is the system the host was built for — read before
anything is applied, so the failure is safe and total. Nothing reports it:
the unit is active, the bus link is up, and the log says it is hearing
what the node should be.

Worse, the declaration that would fix it is the declaration it cannot
apply, so the mesh cannot repair such a machine. Restored by moving the
delivered versions aside and letting the launcher fall back, which is the
fallback working as designed.

0142 already settles the version — it comes from where the component sits,
not from its linker — and that is unimplemented. The system pin has no
answer, and the candidates are a decision rather than a fix: put it in the
path too, carry it in a file beside the binary, or stop pinning at link
time at all, which is 0005's to change.
2026-09-30 12:28:01 +02:00
jschoubben 4f9dc3906e Merge pull request 'A machine says little about itself, and only when the mesh asks it something' (#209) from issue/160-what-a-machine-says-about-itself into main 2026-09-30 08:36:48 +00:00
jschoubben a2542e51f8 A machine says little about itself, and only when the mesh asks it something
Filed as a to-do. Nothing is broken by it: every machine here is amd64 and
reports so, and one architecture is enough for now.

When a machine joins, the mesh should collect what it reasonably can about
it and refresh that daily. It already asks what a machine can do; what it
is made of is the same question one level down.

More is already collected than it looks — eight capabilities, the links
that face outside, the host version, and on an adopted machine what it
holds, what is reachable and the firewall and tunnel it was found with.
The architecture and the kernel are in there too, and nothing reads
either: measured, all four machines report amd64 and linux, and node show
prints the capabilities beside them without printing them.

Missing: memory, disk, the processor beyond its architecture, the
distribution and its version, virtual or physical, cores, uptime. Several
are what somebody wants when deciding where a module goes, and the
placement code's own comment already imagines them.

Also missing: the refresh. A machine publishes after an apply, and the
five-minute reconcile publishes nothing, so the mesh's picture is as old
as the last push. Same mechanism 087 wanted.

This is the third thing in one day found to be collected and read nowhere,
after held resources and the host version. Whatever gets added should say
in the same breath which surface shows it, or it will be the fourth.

159 gains the note that the architecture is already reported, so matching
an artifact to a machine needs no new fact — only the comparison and a
compiler told what to target.
2026-09-30 10:36:40 +02:00
jschoubben b19b29cd3a Merge pull request 'An artifact's system is checked and then nothing uses it' (#208) from issue/159-an-artifacts-system-is-checked-and-ignored into main 2026-09-30 08:30:52 +00:00
jschoubben 4bf4fe2d06 An artifact's system is checked and then nothing uses it
Asked whether the host is built for more than one architecture. It is
not, and the reason is worse than a missing feature.

A bundle in a compiled language must name a system, must name one of
alpine, android or arch, and is refused with a careful message if it gets
that wrong. The field is then read by nothing: it does not reach the
compiler, no machine is matched against it, and nothing chooses between
two artifacts by it. The compile runs with no target named and produces a
binary for whatever the build machine happens to be.

The host is x86-64 because the build machine is, not because the
declaration said so. Correct for this mesh by coincidence — four machines,
all x86-64 Arch.

A module declaring two systems would get two identical binaries, both
published and both pinned, and the one sent to the machine it was not
built for would fail at exec. android is the sharp end: not an x86-64
platform, and an artifact declared for it today would be an x86-64 binary
wearing the label.

A field that is checked and ignored is worse than one that does not
exist, because the check is what persuades you it works.

Also noted: the processor is a second dimension the manifest has no word
for, so even implementing the present field would not answer the question
that found this. And since the Go toolchain builds statically, one binary
would run on all three systems anyway — so the pin is a policy rather
than a necessity, which is a decision and not a fix.
2026-09-30 10:30:45 +02:00
jschoubben 5e12081774 Merge pull request 'Issue 142: the mesh compiles its own host and publishes it to its registry' (#207) from issue/142-the-mesh-can-build-its-own-host into main 2026-09-30 08:05:12 +00:00
jschoubben 9ad64881a9 Issue 142: the mesh compiles its own host and publishes it to its registry
Both of the things ADR 0141's insight named as remaining are built. A Go
toolchain based on a new mesh-tools-go module, so the compiler is named
and not pinned; and ${version} in any value of a resource that uses an
archive or a bundle.

Measured rather than asserted: the mesh built the host through its own
toolchain, published it to its own registry, and the bundle fetched back
out is a statically linked stripped binary that runs and says it is the
host.

The cost was larger again than 0141's note said. Three more things in the
path assumed one language or one shape — an entrypoint became a .ts file
whatever the language, the output directory was the compiler's to create,
and a bundle was refused if it named what it is built from — and a fourth
was in the base image, which is Alpine where the first Dockerfile ran
apt-get. That last one is issue 136 in an image, and the build refused
rather than a module failing later.

The version in a path is the digest, not the commit: two builds of one
commit are the same bytes, so a content-addressed version keeps the path
an unchanged build already had.

Still nothing delivers a version to a machine. The host module declares no
resources, so the bundle sits in the registry and no machine is asked to
take it. 0141 carries the insight and 142 the account.
2026-09-30 10:05:04 +02:00
jschoubben 266ade6e28 Merge pull request 'Issue 087: what I shipped first said the opposite of the truth' (#206) from issue/087-a-commit-has-no-order into main 2026-09-30 07:29:53 +00:00
jschoubben a794ef9a3f Issue 087: what I shipped first said the opposite of the truth
It reported "N machines run an older host than another" by comparing
versions as strings. A host reports its version as a commit, and commits
have no order. On the live mesh it named the three machines running the
NEWER host as the ones behind — ced54d4 sorts above 04a27ca and that is
all it means.

The code even carried a caveat saying versions compare as strings and that
this "is enough for the timestamps and commits this mesh uses". That was
the error, written down and not noticed: enough for timestamps, meaningless
for commits, and the mesh reports commits.

It now reports the split and claims no ordering, which is more useful as
well as more honest — the reader sees who is on which side, and that is
what decides whether a field can be sent. Ordering is left with the host,
which would have to report something ordered for anybody to have it.

This is issue 145 arriving by my own door an hour after I closed it: a
report that confidently says the opposite of the truth is worse than one
that says less.
2026-09-30 09:29:40 +02:00
jschoubben e10084fed6 Merge pull request 'Issue 087: a machine states its host only when the mesh sends it something' (#205) from issue/087-a-machine-says-its-host-only-when-asked into main 2026-09-30 07:09:17 +00:00
jschoubben 941b920bd7 Issue 087: a machine states its host only when the mesh sends it something
Measured live. All four machines run the identical host binary — same
digest, installed within eighteen seconds — and at first only one reported
a version, which read as a difference where there was none.

A machine publishes a report after an apply. The five-minute reconcile
publishes nothing, because it is the machine keeping itself as declared
rather than answering anything. So a current, idle machine never says, and
the mesh cannot tell that from a machine running something ancient.
Confirmed by pushing: not reported, then 04a27ca.

Enough for the purpose, not enough for the claim. For deciding whether a
new declaration field is safe it is sufficient — pushing is what the mesh
is about to do, and the answer arrives with the act. For knowing what the
mesh runs it is not, and "not reported" is worded as "nobody has asked
recently" for that reason.

Making a heartbeat carry it would close the gap and would change what a
heartbeat is — a bare word that the node is there, deliberately carrying
nothing else. Left alone rather than widened in passing.
2026-09-30 09:09:10 +02:00
jschoubben ecbd2ce2e6 Merge pull request 'Group 1: 145's report states its scope, and 107 waits for delivery' (#204) from issue/145-and-107-what-group-one-leaves into main 2026-09-30 07:00:32 +00:00
jschoubben b7f7b97d8a Group 1: 145's report states its scope, and 107 waits for delivery
145, partly resolved. The sentence that was true for eleven hours of a
mesh in which no module could reach another now says what it is not a
claim about: that is the mesh and the machines agreeing, and nothing here
dials a provision. It checks nothing and does not pretend to — ADR 0146
decides the check and is deliberately not built. What changed is that the
report no longer implies otherwise. Stays open for that reason.

Carried forward: 0146's check needs an internal name fetched over TLS with
the certificate verified, and until today no machine trusted the mesh's
authority. Three of four do now, so whoever builds it does not have to
solve that first.

107, diagnosed and deliberately not built. The premise is confirmed in the
host's own words — unknown fields are refused because "a field the host
does not know is a thing the control plane believes it asked for" — so the
fix is a flag day, not an addition. 087 now makes the cost measurable, and
the measurement is why it waits: one machine of four runs an older host,
it is parked, and nothing delivers a host at all (142). Shipping the field
means hand-placing binaries and unparking a machine, and one missed in
that sequence is unreachable, not degraded. The fault it prevents has
never been observed.

142 gains the note that it is 107's gate, and that it is what makes a
declaration field cost a rollout instead of an expedition.

A judgement about order, not a refusal, and cheap to overrule.
2026-09-30 09:00:18 +02:00
jschoubben c2cf0d72d6 Merge pull request 'Issue 087 is resolved: the mesh knows which host runs a machine' (#203) from issue/087-the-mesh-knows-which-host-runs-a-machine into main 2026-09-30 06:55:05 +00:00
jschoubben 0a5006b366 Issue 087 is resolved: the mesh knows which host runs a machine
The machine has reported its host version since ADR 0141, whose own
comment says why it must: without it nothing can say a machine is behind.
The controller's copy of the report did not have the field, so it
unmarshalled into nothing and was thrown away on arrival. Two structs
describe one message and only the sending side had it.

node show names it per machine, "not reported" where the mesh has not been
told. status names every machine running an older host than another does,
and which is newest.

Disagreement rather than staleness, deliberately: nothing delivers a host
version yet, so the mesh holds no canonical current one and "behind" has
no fixed point. What it can say is that the oldest host in the mesh is
what the mesh may send.

Two refusals to guess: a machine that reported nothing is not called
behind, and versions compare as strings — right for the timestamps this
mesh uses, wrong for a scheme where 10 sorts before 9, said at the place
that would have to learn.

107 gains the note that this is what makes its new field safe to consider,
and that one machine of four is behind today, so it is not free yet.
2026-09-30 08:54:52 +02:00
jschoubben 3ea9601904 Merge pull request 'Issue 125 is resolved: a hold is a line in the report, and it breaks "all well"' (#202) from issue/125-a-hold-is-a-line-in-the-report into main 2026-09-30 06:47:57 +00:00
jschoubben f7a37ee4f3 Issue 125 is resolved: a hold is a line in the report, and it breaks "all well"
Two of the four surfaces the report named already carried it — the host
has reported Held since ADR 0100, and node show reads the machine's own
list with an `as of` beside it. Recorded as checked rather than assumed.

Two did not. The apply line counted what it applied and said nothing
about the difference; status read the mesh's take-time listing, so a
module assigned after it showed nothing at all.

Both now say it, and the part that carries the weight: a hold suppresses
"all doing what they were told, all heard from, running what the mesh
would send them". That sentence was true for the whole outage, and acting
on it is what stopped the predecessor's proxy. Being adopted still does
not suppress it — a mode somebody chose is not a half-finished action.

Status does not call a hold a fault, deliberately. It is correct
behaviour, and a reader trained to see red for something the mesh did
right stops reading.
2026-09-30 08:47:36 +02:00
jschoubben 601d004fcb Merge pull request 'Issue 129: three machines of four trust the mesh, not one' (#201) from issue/129-three-machines-not-one into main 2026-09-30 00:30:30 +00:00
jschoubben d009c3efef Issue 129: three machines of four trust the mesh, not one
Extended after the first was proven. Every converged machine now holds
the anchor and verifies an internal name with a plain client; before,
the two unassigned ones answered 'unable to get local issuer
certificate' and held no entry for the mesh.

ace is excluded on purpose: it is adopted, so a module assigned there is
held rather than run, which is right and is not trust.

Both the resolution and 0147's insight said one machine of four, which
was true for about twenty minutes.
2026-09-30 02:30:23 +02:00
jschoubben 5036b927b9 Merge pull request 'Issue 129 is resolved: a workstation trusts the mesh, and stops when told to' (#200) from issue/129-a-machine-trusts-the-mesh into main 2026-09-30 00:18:51 +00:00
jschoubben 1f72e82b84 Issue 129 is resolved: a workstation trusts the mesh, and stops when told to
Registered ca-trust from the catalogue — it was merged and had never been
registered, which is why "assign it to one machine" had no module to name
— assigned it to the workstation, and verified.

Verified in the form ADR 0147 prescribes, against the authority's own API
so the handshake needs nothing else in the mesh to be right: 200, issuer
Mesh Internal CA, Verify return code 0. Four routed internal names verify
too, and `git ls-remote https://…` works, which is the consequence the
report named.

Removal exercised for the first time. Unassign and push removes the
anchor, empties the trust store of the mesh's authority, and returns the
plain client to the original error; assigning again restores it. That is
the half 0147 claimed and nothing had shown.

One thing it found that is not in the module: removal works only because
the host removes the service before the script. Stopping the unit is what
deletes the certificate and refreshes the bundles, and it needs the script
to still exist. The symmetry rests on an ordering nothing states.

0147's "written, and not yet run" now carries a progressive insight
saying it has run, where, and that it ran on one machine of four.
2026-09-30 02:18:21 +02:00
jschoubben 76fbe323ea Merge pull request 'Issue 118 is resolved: it was issue 135, and umami is healthy' (#199) from issue/118-is-135-and-is-resolved into main 2026-09-29 23:54:39 +00:00
jschoubben 295bf2f3ab Issue 118 is resolved: it was issue 135, and umami is healthy
Verified on the machine: umami has zero restarts, applies all 26
migrations, reports the store up to date, and umami.novox.be answers 200
where it had answered 502 since 2026-09-25.

It was the same fault as 135, filed three days earlier and diagnosed
without either record noticing the other. 135's container IS umami — it
is named in 135's own evidence table, holding novox.internal:10.42.0.1
after the overlay range moved. The dial appeared to succeed and the first
real query timed out because the name pointed at an address that no
longer existed; the store, the path and the credential were all fine.

118's own reasoning is kept as a warning, because it is careful and
wrong: it argued path-MTU and conntrack, and it named the move that would
have found it — compare umami against a container created that week —
and did not make it. A stale name presents as a network fault, which is
why ADR 0148 stops copying names into containers rather than detecting
when the copies go bad.

Also: 118's located-in named mesh-catalog modules/umami, which was never
at fault; it names mesh-host's comparison now. And 135 gains the pointer
to 118, which is the back-reference I have now missed three times.
2026-09-30 01:54:19 +02:00
jschoubben e1b74810a8 Merge pull request 'Issue 129 is live and reproduced, and needs three steps rather than one' (#198) from issue/129-and-what-reproducing-it-found into main 2026-09-29 23:30:48 +00:00
jschoubben e41eed0852 Issue 129 is live and reproduced, and needs three steps rather than one
The certificate is genuine, from Mesh Internal CA, and nothing on the
workstation trusts it — verbatim the error the report gives. The public
name on the same proxy verifies cleanly, which puts the fault exactly
where the report puts it.

What is in the way is not an assignment. `ca-trust` is merged in the
catalogue and has never been registered with the mesh — 39 of 76
manifests are — so there is no module to assign. It dry-runs clean and
needs no artifact built.

Two findings from reproducing it, both their own issues:

157 — every routed name is published with an `.internal` alias that
nothing serves. The hosts file says keycloak.novox.be.internal; the proxy
serves keycloak.novox.internal and refuses the other by name. The first
three names I tried came from the hosts file and failed with a TLS alert
rather than a verification error, which pointed at a regression that had
not happened.

158 — the proxy re-logs all 52 routes every two seconds, 31 times a
minute. The one line that explained 157 sat between two of them.

Also recorded, because it was nearly filed as a defect and is not one:
step-ca publishes roots as /roots.pem, which is PEM, so ca-trust's fetch
and its refuse-a-non-certificate guard are both right. Its other endpoint
/roots returns JSON that contains the text the guard greps for, so the
guard is sound only because of which path is published.
2026-09-30 01:30:24 +02:00
jschoubben ba30286896 Merge pull request 'The pointers back from what yesterday's records changed, which I missed twice' (#197) from decision/the-pointers-back-from-what-these-narrow into main 2026-09-29 22:47:45 +00:00
jschoubben 53b94c51bb The pointers back from what yesterday's records changed, which I missed twice
Three records were left describing a mechanism a new record had moved,
and a reader arrives at them by following a citation: 0066 still said a
routed name is written into every container after 0148 replaced that with
resolution; 0016 still read as though the lab were the test bed after
0149; and issues 109 and 135 said nothing about 0148 ending the copying
that 135's own fix made comparable. Each was a citation leading to the
wrong answer in a record that was not wrong about anything it decided.

This is the second time in one session. The playbook rule I added last
round did not stop it, so the convention is now written where the record
conventions live, with the shape to use and three worked examples — and
with the honest note that it is NOT machine-checked and cannot be from
`extends:` alone: 102 records extend another, 87 have no back-reference,
and that is correct, because extending usually means building on a
context. Making it mechanical means a record declaring the relationship
in frontmatter, which is a schema change and is not mine to decide.

Also: designs 18 and 20 claimed `updated:` dates from before I edited
them, and 117's `fixed-by` gained the commit beside the record.
2026-09-30 00:47:19 +02:00
jschoubben 83791f0921 Merge pull request 'The four open design questions, answered: ADRs 0148, 0149, 0150, and 0114 accepted' (#196) from decision/0148-the-meshs-names-are-resolved-not-copied into main 2026-09-29 22:39:38 +00:00
jschoubben bf4d4e2e7b The three parked questions are answered: 0114 accepted, 0068 superseded, 117 settled
**0114 accepted.** The separation it draws — a consumer's resource and a
credential that reaches it are different things — is a data-loss rule,
and a record naming one should not sit unresolved while the code that
could hit it is written. Not built, and accepting it schedules nothing;
accepted-and-not-built is where 0141 and 0142 already are.

**0068 superseded by 0149, the live mesh is the test bed.** Never built,
and contradicted by practice that was written down nowhere but a handoff
note. The faults that cost the most are faults of a mesh that already
exists — bound consumers, containers made against an older roster, an
adopted machine — and a bed is by construction a mesh that does not.
Issue 156 settles it: that change was verified honestly, on the only path
where it cannot fail. The lab is not retired; raising a mesh from bare is
now its whole job.

**117 answered by 0150.** A module's own code runs as supervised
processes under the module's one account. 0047's argument was for a
runtime per module, not for a container — a unit satisfies it and runs as
an account. The invariant is the account, not the process count: its
worry was a second identity to scope and seal, and processes sharing one
account create none. Designs 18 and 20 now cite it and 0047 carries a
dated pointer, which closes the third disagreement — that neither design
knew the record existed.
2026-09-30 00:39:09 +02:00
jschoubben ec42ee0846 ADR 0148: the mesh's names are resolved, not copied into every container
Answers issue 151. Copying the roster into each container made the roster
part of each container's identity, so one name moving replaced every
container in the mesh — and it never stopped the staleness it was for,
since a copy taken at creation is stale the moment the roster moves (109,
135).

A container resolves through its machine's resolver instead, and nothing
is copied. Staleness stops being possible rather than detected, and a
name's blast radius becomes nothing.

Scoping each container to the names it binds was the close call and is
rejected: it contradicts anything-calls-anything, and leaves the roster
in the digest so the churn returns for a widely-bound name.

Gated on issue 110 — a container on the runtime's default network has no
DNS at all today. Removing the copy first reintroduces 109 and 135
silently on a live mesh. 151 stays open until the code lands; design 08's
file-not-resolver passage is narrowed to the machine's own roster.
2026-09-30 00:34:30 +02:00
jschoubben 5a3dee9e9e Merge pull request 'The records pointed at branches that no longer exist, and two fixes had no sequel' (#195) from issue/pointers-that-resolve-to-nothing into main 2026-09-29 22:29:02 +00:00
jschoubben 47909c6b71 The records pointed at branches that no longer exist, and two fixes had no sequel
Three issues named the branch that fixed them, and a branch is deleted
when it merges — so every `fixed-by:` was a pointer that resolved to
nothing by the time anyone followed it. They name commits and pull
requests now, and playbook 03 says to.

Two records were missing the thing a reader arrives for. 146 did not say
that one of its fixes crash-looped the control plane on a running mesh,
which is the whole reason the delivery subject carries the stream and the
raise path was the only one exercised. 151 did not say that 152 removed
the false reasons its roster moved, or that it stays open for the real
ones.

ADR 0080 enumerates what cycle.py enforces and named four things; it
enforces five. A progressive insight names the fifth — the decision
stands, the list had gone stale. The checks README and playbook 03 gained
the same rule, and 155 points at all three.
2026-09-30 00:28:35 +02:00
jschoubben 7e5edab8da Merge pull request 'Issue 156: the wider grant has landed, and the notes the fix printed' (#194) from issue/156-the-grant-has-since-landed into main 2026-09-29 22:19:35 +00:00
jschoubben 3649f82204 Issue 156: the wider grant has landed, and the notes it printed
Measured on the broker's own config after the fixed controller composed
it: every node is now allowed _DELIVER.<node>.> as well as the bare
subject. That grant was missing when this was diagnosed, which is why
re-making the consumers would have silenced every machine.

Also records the five consumers it kept and reported, including the build
machine's worker, which is bound the same way and was not anticipated.
2026-09-30 00:18:57 +02:00
jschoubben b400ce2c54 Merge pull request 'Issue 156: moving a consumer's delivery subject stops a running mesh' (#193) from issue/156-a-consumer-that-works-is-not-replaced into main 2026-09-29 21:57:00 +00:00
jschoubben 995c8cb266 Issue 156: moving a consumer's delivery subject stops a running mesh
146's change is correct on a mesh being raised and fatal on one that is
running: the server will not move a push consumer's delivery subject
while a subscriber is bound, and a node is bound to its declaration
consumer the whole time it is up. Reproduced before fixing; an earlier
version of the test unsubscribed first and passed against the broken
code.

The consumer that works is kept and the assertion says so. Not re-made:
the nodes are granted the bare subject only, so re-making would have
silenced every machine.
2026-09-29 23:56:31 +02:00
jschoubben f89aef992d Merge pull request 'Two records shared a number, twice, and every check passed' (#192) from issue/two-records-share-a-number-and-nothing-says-so into main 2026-09-29 21:39:09 +00:00
jschoubben b967ef7be3 Two records shared a number, twice, and every check passed
Two machines filing issues in the same hour both read main correctly and
both took "the next free number". main lags every open pull request —
seven that evening — so they collided twice. The second collision
reached main with records, cycle and index all reporting success.

cycle.py now refuses a tree where two issue folders share a leading
number, and names both. Proven by adding a duplicate and watching it
fail. The colliding records become 153 and 154, renumbered in the branch
that lands last, because renumbering a branch whose author is still
pushing only moves the race.

The check catches a collision; it does not prevent one. Taking a number
still means reading the open pull requests as well as main — issue 155
says so.
2026-09-29 23:38:48 +02:00
jschoubben e417906241 Merge pull request 'Issue 150: a machine's own network is not a reach' (#190) from issue/150-a-machines-own-network-is-not-a-reach into main 2026-09-29 21:37:41 +00:00
jschoubben b1bf895688 Merge pull request 'Issue 149: an adopted machine's data cannot be placed where it is' (#189) from issue/149-adopted-data-cannot-be-placed-where-it-is into main 2026-09-29 21:37:34 +00:00
jschoubben 0d64677c70 Merge pull request 'Issue 152: a node the mesh could not read withdrew its names from every machine' (#191) from fix/152-a-lookup-failure-is-not-an-absence into main 2026-09-29 21:37:16 +00:00
jschoubben 04205c1dc8 Merge pull request 'Issues 147 and 148: a route before its module is taken; a new name recreates every container' (#188) from issue/147-148-found-migrating-ace into main 2026-09-29 21:37:09 +00:00
jschoubben 9dc49cd831 Merge pull request 'Grooming: five issues were fixed and never closed, and one is not' (#187) from grooming/stale-issues into main 2026-09-29 21:37:04 +00:00
jschoubben 66b413a076 Merge pull request 'Issue 140 is resolved, and was resolved before it was read again' (#186) from issue/140-resolved into main 2026-09-29 21:36:57 +00:00
jschoubben fcba05fed9 Merge pull request 'Issue 146: a first node now enrols, and is enrolled twice' (#185) from issue/146-diagnosis into main 2026-09-29 21:36:42 +00:00
jschoubben 18f37c25b2 Issue 152: the loop is metastable, and it cleared at 23:11
It ran 22:46-23:11, five full replacements, and stopped when a pass
happened to read the store during a window it was up. The record said
nothing about it stops; that was wrong. Exiting by luck is the finding,
not a mitigation.
2026-09-29 23:27:51 +02:00
jschoubben 3c2b4fc6b6 Issue 152 is fixed: a lookup failure is no longer an absence
The three gatherers pass over a node whose set does not compose, and
raise anything else. 151 stays open: this removes the false reasons a
roster changes, not the fact that a real change still replaces every
container.
2026-09-29 23:26:20 +02:00
jschoubben 72eaf52867 Issue 152: a node whose plan will not compose withdraws its names from every machine
ace numbered its two records 147 and 148, which this repository already
uses; they become 150 and 151, as 127 became 149.

The new record is why the control node could not stop applying: the
roster alternates between two values because routeNamesInTheMesh
swallows a per-node plan failure, and the roster is part of every
container's identity. The loop closes through the control plane's own
store, which each pass replaces.
2026-09-29 23:13:54 +02:00
jschoubben 741625e725 Merge remote-tracking branch 'origin/issue/147-148-found-migrating-ace' into issue/the-roster-flicker-recreates-every-container 2026-09-29 23:12:57 +02:00
jschoubben c3730b9a23 Issue 150: a machine's own network is not a reach
ADR 0138's reach is internal, public or both. ace's LAN devices (an IoT
switch on mosquitto, unifi access points, plex clients) are none of them;
the only reach that admits them is public, which states the wrong thing and
opens the port to the internet on any machine with a public address.
2026-09-29 23:06:08 +02:00
jschoubben 3a59099c81 Issue 149: an adopted machine's data cannot be placed where it is
ADR 0112 decided that an assignment places a directory and says where an
access's data is; the control plane implements neither. ace's media stack
(a 40 TB ZFS library, plex state on a second disk) cannot be expressed
without writing ace's paths into manifests.
2026-09-29 23:03:11 +02:00
jschoubben 3bd6f34de3 Issues 147 and 148: a route before its module is taken; a new name recreates every container
Both found migrating the first module on ace (searxng), adopted with route-adapter:
147 — assign is not a no-op for a routed module: its route reaches the predecessor's
proxy while its container is held, and the name serves a dead backend until take.
148 — every push that moved the mesh's names replaced every container on novox, twice,
the control plane's own store included.
2026-09-29 23:01:51 +02:00
jschoubben 2b5119ecd2 Issue 114 is answered: the controller is a process, by ADR 0142
It asked a question rather than reporting a defect, and the question was taken
five days later — the mesh's own components are binaries on the machine, and
third-party software stays a container because an image is the right way to
carry somebody else's build. The delivery of them is issue 142.
2026-09-29 22:38:11 +02:00
jschoubben 96bdffa9bc Two records were numbered 127; the second becomes 149, and is resolved
A number identifies a record, and two were given 127 on 2026-09-27. The one
three documents and three source files cite by number keeps it; the other
becomes 149, says so in its own heading, and its two inbound references are
repointed.

It is also resolved: an empty declaration is sent carrying owns_nothing rather
than skipped, and the host refuses an empty body that does not carry it, so
emptiness cannot be read as a truncated declaration.
2026-09-29 22:33:40 +02:00
jschoubben 4eb16f1028 Three proposed records were already built; two are still yours to call
0037 (where a module lives), 0113 (the vault makes every secret) and 0115 (one
assignment of a module per node) described arrangements the mesh has — a
catalogue of other people's software plus a table filled by module add, a vault
answering a secret requirement six modules make, and a rule the assignment
table's primary key already enforces. Each is accepted against what was built,
and says so in its own words.

0037's other half is not built: a manifest outside this catalogue has no check,
which is issue 148.

0068 (the lab takes requests) and 0114 (a shared credential rotates over two
credentials) stay proposed. Neither is built, and both are decisions rather
than records of something that happened.
2026-09-29 22:31:55 +02:00
jschoubben d199de40db Merge the 146/147 records, which 006's note links to 2026-09-29 22:31:46 +02:00
jschoubben 9a1dc4665c Grooming: issue 006's knowledge base is the predecessor's, and is gone
The record catches README claiming this repository is indexed into a knowledge
base. Nothing indexed it, and since the cut-over there is nothing to index it
into — the surface that answered is on the transport the mesh removed (issue
147). Noted where the record is, so the next reader does not go looking for a
search that cannot exist.
2026-09-29 22:26:13 +02:00
jschoubben 14be8576f8 Grooming: five issues were fixed and never closed, and one is not
088, 089, 120, 128 and 130 each name a commit that is on main and cites them —
the forge's address following a moved port, a route naming its endpoint, a
provisioner asking the backend what is there, the hosts file written into a
marked block, and undeclaring giving a unit back the state it was found in.
Each says it was closed by reading commits rather than by a run, so nobody
reads a green that was never measured.

129 stays located on purpose: ca-trust is merged and no machine holds it, so
the symptom it opened on is still true everywhere.
2026-09-29 22:25:25 +02:00
jschoubben ec8676c225 Issue 140 is resolved, and was resolved before it was read again
Written 2026-09-28, answered the same week by mesh-controller bdf965d and
c68d3a7, and never closed — so the mesh's own account said reach was declared
nowhere while three readers were reading it: the filter, the proxy's names, and
each authority's host policy. The resolution names them and what checks each.

What is left is retiring the older per-port keys the block replaces, which is
not a gap in what reach can say.
2026-09-29 22:19:24 +02:00
jschoubben 3c0f7082e6 Issue 147: the tool surface is not the mesh's, it is the predecessor's
Diagnosed from the configuration: the tool server is HAL's brain, a local
process on the workstation with the predecessor's broker URL in the assistant's
own config. No manifest, no assignment, no seat, no account. Nothing regressed
— the mesh removed a transport this program still dials, and the program was
never part of the mesh. The mesh has a tool model and nothing publishes an
operator-facing surface onto it.
2026-09-29 21:50:50 +02:00
jschoubben 0dd00e88b6 Issue 147: the operator's tools still dial the bus that was removed
Every tool call fails with 'AMQP not connected', on every node including the
local one, because the tool surface still opens an AMQP connection and that
transport was deleted at the cut-over. The mesh reports healthy throughout —
what broke is the thing standing outside asking it questions, so nothing the
mesh checks is about it.
2026-09-29 21:39:55 +02:00
jschoubben eef54917ec Issue 146: the double enrolment was two consumers sharing a delivery subject
Not about enrolment. A push consumer delivers onto an ordinary subject and
everything subscribed to it gets a copy; the controller's two consumers were
both named after it, so both were given the same subject and the one process
acted on every message twice. Enrolment is where it drew blood because a second
enrolment mints a second credential.
2026-09-29 21:32:52 +02:00
jschoubben e9b1010bc0 Issue 146: what made it slow, and what was changed so it is not 2026-09-29 17:45:25 +02:00
jschoubben f6ed3545b7 Issue 146: a first node now enrols, and is enrolled twice
Two more faults behind the three already fixed. The account a token is the
password of was never recorded, and the comment above the issuing code said
it was; issuing now records it. Placing the composed list at genesis is the
other half — the control plane says what it composed and whoever raises the
machine writes it beside the bus, because no declaration can reach a machine
that has not enrolled.

With that a first node enrols. It is then enrolled twice from one attempt,
each minting a credential, and it keeps the answer to the first while the mesh
keeps the second. The trail for that one stops at a duplicate that survived
message-id deduplication.
2026-09-29 17:36:40 +02:00
jschoubben 0b08cdfce1 Merge pull request 'ADR 0147: a module anchors the mesh's authority, and issue 146: the foundation cannot be raised' (#184) from decision/0147-a-module-anchors-the-meshs-authority into main 2026-09-29 14:06:41 +00:00
jschoubben bffd2af40c Issue 146 diagnosed: four faults stacked, three fixed, the fourth is genesis
Raising a first node hits them in order: the bundle's bus image named for a
registry that is gone; the certificate made by openssl in an image that has
none; enrolment dialling TLS at a bus that speaks first, then refusing its own
token for the empty bus. Each was right until the bus changed and nothing has
raised a foundation since.

The fourth is not a patch: the installer carries the bus's first user list and
the controller composes the rest through the bus being a module, and at genesis
there is no module — so the first node cannot be let onto the bus it just
raised.
2026-09-29 15:42:32 +02:00
jschoubben 4a51ea4b3a Issue 146: the foundation cannot be raised on the bus the mesh runs on
Found trying to run ADR 0147's bed. The older bundle raises the previous
broker and a control plane that refuses to start without MESH_BUS_NATS; the
newer one stops a step earlier, asking for a certificate from openssl in an
image that has none. No bed can run while this holds, so 0147's check section
now says what actually stands behind it — the rendering, not a machine.
2026-09-29 15:25:58 +02:00
jschoubben 6c14d313b8 ADR 0147: a module anchors the mesh's authority on a machine, and takes it away again
Issue 129: every internal HTTPS name fails verification on every machine,
because nothing has ever written the mesh's root into a trust store. The
report proposed the controller inject it the way the private network writes
the registry's trust; this record rejects that — reachability and trust are
not the same fact, and where anchors live is the host's difference, not the
controller's. A module requiring internal-acme-ca does the whole of it, and
being unassigned undoes it.
2026-09-29 15:02:37 +02:00
mesh-admin ced547dae9 Merge pull request 'ADR 0146: connectivity is checked by name, per hosting form' (#183) from decision/0146-connectivity-by-name into main 2026-09-29 12:01:57 +00:00
jschoubben 38482435af ADR 0146: connectivity is checked by name, per hosting form, with a valid certificate
0145's core stands — a module checks what the mesh claims, from where the callers
are — and everything it said about what to dial was wrong.

A raw port is not how anything in this mesh is reached. A real caller resolves a
name, the proxy answers, the proxy reaches the service; dialling a port tests the
last hop of a four-hop path and skips the three where most connectivity lives.

So each hosting form gets its own endpoint, route and name:
connect-docker.<node>.internal for a container, connect-process for the mesh's own
code in a unit it writes, connect-unit for a unit a package ships — and the same
under each public domain. Every machine checks every machine, by name, over TLS,
verifying the certificate against the authority that should have issued it. The
names are the instrument: a failure reads as connect-docker.g14.internal did not
answer.

No name is written anywhere: the machines come from the roster fact, the labels are
the module's, and a machine that joins is checked by the others on the next push.

Five things this needs that the mesh does not have, named rather than assumed —
a module every machine has (ScopeNode is exclusivity, not obligation), a container
running a bundle, a unit a package ships, the public domain in the roster fact, and
issue 129, until which every internal name will correctly fail verification.
2026-09-29 14:01:55 +02:00
mesh-admin cf8a8d78c9 Merge pull request 'ADR 0145: a module checks what the mesh claims is reachable' (#182) from decision/0145-a-module-checks-what-the-mesh-claims into main 2026-09-29 11:44:58 +00:00
jschoubben 9de25994e9 ADR 0145: a module checks what the mesh claims is reachable, and it checks itself
The mesh asserts three things are callable and has never checked any of them. A
module on every machine serves an endpoint of its own and dials every other
machine's, from the position the callers are in — not the host and not the control
plane, both of which reach those addresses by paths no ordinary caller uses and
would have passed throughout the outage.

Its probe is its own endpoint, declared reachable over the private network, so it
is admitted by exactly the rule that governs every internally-exposed service. The
tempting target is a service every machine has, and those are the ones never closed
— ssh above all — which would have passed for all eleven hours.

Resolution and connection reported separately, because they have different owners.
One failure is not a fault, and the count travels with the result. It reports and
repairs nothing.

Built on a scheduled container, a rendered roster fact and a bundle: no new
vocabulary. Adopted on its own merits, which is what 0143 was not.
2026-09-29 13:44:56 +02:00
mesh-admin 64ea47b11d Merge pull request 'ADR 0144: anything on a machine may call anything on it' (#181) from decision/0144-local-is-not-a-boundary into main 2026-09-29 11:32:03 +00:00
jschoubben eba24a72af ADR 0144: anything on a machine may call anything on it, superseding 0143
Everything should be able to call what runs on the same machine, another machine's
service exposed to the private network, and another machine's service exposed
publicly. Three cases; the filter had two.

The first was expressed as the machines' own addresses on the private network. A
caller on the machine carries such an address; a caller in one of its containers
carries a bridge address and matched nothing — measured, same destination and same
machine: src 10.10.0.1 against src 172.17.0.8. The second case worked by accident,
because the tunnel rewrites a caller's address to the sending machine's. Two of
three working is why it read as correct.

0143 answered the wrong question. It proposed verifying each grant from the
consumer's own network position and went to length about which position, because
whether a caller sat in a container changed the answer — and that difference was the
bug. Observing a configuration error is not its remedy. Superseded, and nothing
replaces it; whether the mesh should check a grant is still open in issue 145 and
must stand on its own.

And a module is not a container: 61 of 72 happen to use one, 11 do not, and a rule
reasoning about containers describes most of the mesh rather than the mesh.
2026-09-29 13:32:01 +02:00
mesh-admin 78d4873f4f Merge pull request 'ADR 0143: a consumer verifies the grant it is given' (#180) from decision/0143-a-consumer-verifies-its-grant into main 2026-09-29 11:07:20 +00:00
jschoubben ddfd62edf6 ADR 0143: a consumer verifies the grant it is given
A grant is four facts and a credential — the provision, the machine, the port, and
who the consumer is when it connects — and it is the whole mechanism by which
anything in the mesh reaches anything else. The mesh asserted it and never found
out whether it was true. Issue 145 is what that cost: eleven hours of 'every module
current with its source' while a module could not reach its database.

The consumer verifies it, from its own network position. Not the control plane,
which reaches the address by a path no consumer uses. Not the machine, whose own
packets carried a source address the filter admitted while every container's was
refused — so a host-side check would have passed throughout the outage it exists to
catch. That is inference from the rule that was loaded, stated as such.

A connection and nothing more; speaking each provision's protocol would be a second
implementation of every provision. One failure is not news, only consecutive ones,
and the count is reported rather than the last attempt. A consumer that is not
running reads unchecked, which is a different sentence from broken.

What it costs to be wrong is the constraint on all of it: a check that calls a
working provision broken trains a reader to ignore the report.
2026-09-29 13:07:18 +02:00
mesh-admin f65664640a Merge pull request 'Issue 145: a machine reads healthy while its modules cannot reach each other' (#179) from issue/145-healthy-while-broken into main 2026-09-29 10:37:10 +00:00
jschoubben 492ac7be18 Issue 145: a machine reads healthy while its modules cannot reach each other
Converging the control node closed every path by which a module reached another by
the machine's own name, and it ran for eleven hours while the mesh answered 'all
doing what they were told, all heard from, running what the mesh would send them,
and every module current with its source'.

6,154 database failures in one module's log, beginning at the minute of the flip.
The service accepted TCP and never answered HTTP. Every check the mesh makes
passed, because every check the mesh makes is about the relationship between the
mesh and a machine — applied, current, containers running. None asks whether a
module can reach what it requires, though the mesh composes every grant and so
knows exactly who requires what.

The filter fault is fixed. The eleven hours are the measurement, not the bug.

Also records on 144 that 'more closed than the mesh believes' was harmless only
for what is reached from outside, and an outage for what is reached from within.
2026-09-29 12:37:08 +02:00
mesh-admin 5ac77e3cef Merge pull request 'ADR 0138: reach asks for names on a routed endpoint' (#178) from decision/0138-insight-reach-and-the-proxy into main 2026-09-29 00:50:28 +00:00
jschoubben a619022c35 ADR 0138: a progressive insight — reach asks for names on a routed endpoint
The record says internal and public each mean something to the filter. For an
endpoint the proxy serves, the second half is wrong, and ADR 0045 said so first:
a public service is exposed through the proxy, listening from the mesh, not by
opening its own port.

Found by trying to express one real module, not by review — routed name public
because browsers post to it, machine port private because it serves a dashboard in
cleartext. Under one value for both, saying public would have reopened a port an
operator had just closed. Measured the same evening: the routed name answered from
the internet over TLS while the port was refused from the same place.

Corrects a fact. One statement per endpoint with three things derived from it
stands; the filter column applies to an unrouted endpoint.
2026-09-29 02:50:25 +02:00
mesh-admin 49c065c204 Merge pull request 'Issue 143: correct the diagnosis' (#177) from issue/143-corrected-diagnosis into main 2026-09-29 00:33:13 +00:00
jschoubben ddb3980f09 Issue 143: correct the diagnosis — the step exists and did not fire
The first account said the declaration carries no resource that would disable the
found firewall and that nothing implemented the sentence the flip prints. Wrong:
the mechanism is a step in the host's own apply, retireFirewall, and it is careful
— it reads back that the mesh's table is loaded before retiring anything, records
the forward policies first, and verifies ufw reports inactive afterwards.

What is established: ufw was active and enabled two minutes after a flip that
reported the node converged with nothing failed; and the machine's record now
reads disabled_by_mesh: true, written by a reconcile fifty minutes later that
found ufw already inactive because an operator had disabled it by hand.

So the step did not take effect and the record says it did. The candidates are
named rather than chosen — the step is skipped silently when the apply has any
failure, the mesh's table is loaded by a service in the same apply so ordering is
open, and the host's detail lines do not reach the journal, which is why this has
candidates instead of a cause.
2026-09-29 02:33:10 +02:00
mesh-admin 75a8f6abc7 Merge pull request 'Issues 143 and 144: the found firewall is neither retired nor all of it' (#176) from issue/143-and-144-the-found-firewall into main 2026-09-28 23:36:10 +00:00
jschoubben 862253518f Issues 143 and 144: the found firewall is neither retired nor all of it
Both found by converging the control node — the first machine with a firewall to
flip, since the two before it had none.

143: the preview and the flip both say the found firewall is disabled. The node
reported converged, 372 resources applied, nothing failed, and ufw is still
enabled and active. A converged node's declaration carries no resource that would
disable it; the sentence is printed by the command and nothing implements it.

144: ufw was never what filtered the traffic that mattered there. Fifty forwarded
openings converged through it had matched zero packets, while a chain the
predecessor installed in the container runtime's pre-accept hook did the work —
in memory only, recreated by nothing. The mesh's filter now covers that path, so
the machine no longer depends on it, but the chain remains and is the only thing
refusing the bus and the registry, which the design requires reachable from
anywhere so a machine can enrol before it has a private address.
2026-09-29 01:36:08 +02:00
mesh-admin 51ef3eb7e2 Merge pull request 'ADR 0142: the mesh delivers its own components as binaries' (#175) from decision/0142-mesh-delivers-its-own-components into main 2026-09-28 22:51:44 +00:00
jschoubben 346e613995 ADR 0142: the mesh delivers its own components as binaries, not container images
Measured: the host is a binary somebody copied to four machines, owned by no
package and built by nothing, while the controller, catalogue, builder and vault
are container images publishing no ports at all. Same language, same project,
same kind of work, delivered two ways — and the difference is not a judgement
about either, it is that images are the only delivery that works.

What it costs: genesis must raise a container runtime before the control plane
can exist; updating the control plane goes through a registry the control plane
runs; a host change cannot be rolled out at all; and compiling the language the
mesh is written in is not a capability of the builder, so the controller is built
from a hand-written Dockerfile — the incantation the bundle toolchain exists to
abolish.

Third-party software stays a container: the store, the registry, the broker are
somebody else's build. The container runtime stays on the machine for modules.
What changes is that the control plane no longer needs it to exist.

The receiving half is already built and tested (ADR 0141). Staged: compile Go, an
artifact names its target, deliver a binary, the host first, then the rest, genesis
last.
2026-09-29 00:51:42 +02:00
mesh-admin 25ca9898d5 Merge pull request 'ADR 0141: a progressive insight on what delivery costs' (#174) from decision/0141-progressive-insight-on-delivery into main 2026-09-28 22:32:17 +00:00
jschoubben 709c095387 ADR 0141: a progressive insight — the delivery is not 'nothing new'
The record claimed a version reaches a machine as an ordinary archive with
nothing new needed. Two things it needs do not exist: no toolchain can compile
the host (the list is typescript and python, and the control plane, also Go, is
built as an image from a Dockerfile instead), and nothing interpolates a built
version into a resource path, so nothing can ask for .../versions/<version>/.

The decision, the options weighed and every consequence stand — the host half is
merged and tested. What was understated was the cost, so it is corrected in place
and dated rather than superseded.
2026-09-29 00:32:14 +02:00
mesh-admin c262de3833 Merge pull request 'ADR 0141: the host delivers its own successor' (#173) from decision/0141-the-host-delivers-its-own-successor into main 2026-09-28 22:22:28 +00:00
jschoubben a815433214 ADR 0141: the host delivers its own successor, and versions live side by side
The supervision was already right — a clean exit means the host stood aside and
the launcher runs what is on disk, failures are counted, and a rollback happens
at the limit. Two things made it dead code: nothing told the running host a
successor was waiting, and the rollback resolved its known-good version through
pacman, which no machine here uses and which two of three operating systems do
not have.

Keeping a version rather than a path was the clue. Versions live side by side in
directories named for them; the newest runs; the running one stands aside between
reconciles; a reconcile that completes records itself and retires what is older
than its predecessor; rollback starts that predecessor. No new resource kind and
nothing new on the bus — an archive already fetches by digest, and the path
written is never the path executing.

Answers issue 142.
2026-09-29 00:22:04 +02:00
mesh-admin 61dc90e508 Merge pull request 'Issue 142: the host is the one thing the mesh does not deliver' (#172) from issue/142-the-host-is-not-delivered into main 2026-09-28 22:03:58 +00:00
jschoubben a7cf5c0c1b Issue 142: the host is the one thing the mesh does not deliver
A host change merged yesterday reached no machine without a person copying a
file. The host is not a build target, no declaration delivers it, and the half
that recovers from a bad host — noticing the executable changed, a known-good
record, a launcher that rolls back — is written, tested and called by nothing.
All four machines run a byte-identical hand-copied binary that no package owns
and no record names, so nothing can say a machine is behind.

Found because ADR 0140 needs the machine to report a new fact, and merging that
could not roll it out.
2026-09-29 00:03:41 +02:00
mesh-admin c6f86ae935 Merge pull request 'ADR 0140: the filter constrains what arrives from outside' (#171) from decision/0140-filter-constrains-what-arrives-from-outside into main 2026-09-28 21:32:10 +00:00
jschoubben 1aeb4fe8d8 ADR 0140: the filter constrains what arrives from outside, and says nothing about a machine's own guests
Reading a converged machine's rendered rules showed the cause: the chain blocks
everything passing through and then allows the machine's own containers back by
listing their address ranges. 0137 made that list typeable and 0139 tried to
generate it; both refined a list that should not exist, because the mesh has no
position on a container reaching outward. Constrain what arrives from outside,
allow what did not, and let the machine report which links face outside — one
fact instead of a list. Ports keep following the modules unchanged.

The records check now allows one record to supersede several, and stops
requiring a withdrawn record's own citations to be live.
2026-09-28 23:31:50 +02:00
mesh-admin 08108b569d Merge pull request 'ADRs 0138 and 0139: an endpoint's reach, and networks forwarded because a module declared them' (#170) from decision/0138-endpoint-reach-and-0139-declared-networks into main 2026-09-28 21:03:49 +00:00
jschoubben 14ff89fa40 ADRs 0138 and 0139: an assignment binds an endpoint and says how far it reaches, and a network is forwarded because a module declared it
Both follow from the same rule the mesh is built on — a node's configuration is
composed from the modules assigned to it. Reach was settled separately by the
filter, the proxy's names and the certificate authority, so "this must not be
public" could not be written; it becomes one value on the assignment that all
three read. And the forward chain consulted two constants plus a typed list
although modules already declare their networks; it now forwards what they
declared, with the host rendering the addresses it allocated.
2026-09-28 23:03:32 +02:00
mesh-admin 2126e7b2cb Merge pull request 'Issues 140 and 141: an endpoint's reach, and a forward chain that does not follow the modules' (#169) from issue/140-endpoint-reach-and-141-forward-chain into main 2026-09-28 20:58:03 +00:00
jschoubben dcdfcf104e Issues 140 and 141: an endpoint's reach is declared nowhere, and the forward chain follows constants instead of the modules
Found preparing the control-node's convergence. Reach is settled independently by
the filter, the proxy's names and the certificate authority, so "this must not be
public" cannot be written and a public certificate is obtained regardless. And the
forward chain allows two hardcoded ranges plus a typed list, though the mesh
already knows which networks exist because its own modules declared them — a range
wide enough to keep four of them would have forwarded two predecessor leftovers too.
2026-09-28 22:57:38 +02:00
mesh-admin d23ace1646 Merge pull request 'Issues 138 and 139' (#168) from issue/138-the-uplink-seat-and-139-an-internal-route-name into main 2026-09-28 19:46:22 +00:00
jschoubben 5dbde0b13a Issues 138 and 139: a seat with interchangeable holders that are not, and an internal route name that resolves to the wrong machine 2026-09-28 21:46:20 +02:00
mesh-admin db3868e2b0 Merge pull request 'ADR 0137: a machine says which networks it routes' (#167) from decision/0137-a-machine-says-which-networks-it-routes into main 2026-09-28 19:31:03 +00:00
jschoubben 05039c4f10 ADR 0137: a machine says which networks it routes, and issue 137 is how that was found 2026-09-28 21:31:01 +02:00
mesh-admin b7bf601ca4 Merge pull request 'Issue 136: a module may name a program the machine does not have' (#166) from issue/136-a-module-may-name-a-program-the-machine-lacks into main 2026-09-28 18:50:37 +00:00
jschoubben bff32e3370 Issue 136: a module may name a program the machine does not have, and everything reports success 2026-09-28 20:50:35 +02:00
mesh-admin d6b62387f2 Merge pull request 'Issue 135: a container's mesh names are not compared' (#165) from issue/135-a-containers-mesh-names into main 2026-09-28 15:19:37 +00:00
jschoubben 4789624857 Issue 135: a container's mesh names are not compared, so a moved address is never noticed
One container restarted 2286 times over five days while the mesh reported the machine as doing what it
was told. Its overlay address was five days out of date: the host compares a container by a digest of
its spec, and the mesh's names were not in it, so a container whose image and files never changed was
left alone holding a name that no longer resolved. Forty-eight others were current only because
something else had recreated them.

The same fault as issue 045, in the field that was left out. Resolved by putting the names in the
digest.
2026-09-28 17:19:35 +02:00
mesh-admin e00862e317 Merge pull request 'ADR 0112 is accepted, and issue 134 records what it is not yet' (#164) from decision/0112-accepted-and-issue-134 into main 2026-09-28 15:12:15 +00:00
jschoubben 3a92e4b80c ADR 0112 is accepted, and issue 134 records what it is not yet
0120 was already accepted; what failed the check was that it rests on 0112, still marked proposed —
and so do four designs. The decision stands: a definition names no node, no mesh and no host path, and
everything a module needs is a requirement the mesh resolves.

Accepting it makes the gap visible rather than hiding it, so issue 134 states it. 0112 says how it is
checked — 'a catalogue test finds no domain name in any definition value' — and there is no such test.
Asked by hand: seven modules name this installation in a value the mesh acts on, and eight mention a
public name in prose nothing reads. The two are not the same fault and the fixes differ, which is why
the issue separates them rather than counting to fifteen.
2026-09-28 17:12:12 +02:00
mesh-admin bfddf78bf3 Merge pull request 'Design 32: what shipped today, and the records it rests on' (#163) from design/28-and-32-what-shipped into main 2026-09-28 14:50:44 +00:00
jschoubben 0b291d89f3 Design 32: what shipped today, and the records it rests on
Two of its statements are built — a version preparing its state, and the mesh saying what it applied —
so the document is in-progress rather than proposed, and names the code that owns them. Resting on a
live record rather than a superseded one: 0127 was replaced by 0131.

What this exposes is pre-existing: it also rests on ADR 0112, which is still proposed, and a document
that is not itself proposed may not. ADR 0120 has rested on it the same way for a while. Accepting or
superseding 0112 is a decision, not a cleanup, so it stays visible in the check rather than papered
over.
2026-09-28 16:50:41 +02:00
mesh-admin a924efcc28 Merge pull request 'ADR 0136: a step gates its module, not the machine' (#162) from decision/0136-a-step-gates-its-module into main 2026-09-28 13:38:18 +00:00
jschoubben b421c73a5a ADR 0136: a step gates its module, not the machine
ADR 0135 made a step something the mesh derives for any module that prepares its state, which turned
ADR 0052's reach into a fault: a module whose database is briefly unreachable would stop every module
declared after it on that machine — the fault issue 011 already removed for every other shape, and
the reason the catalogue migrates itself at start rather than in a step.

A step now stops the rest of its own module and nothing else; an action still gates the machine,
because genesis is a row of them and they belong to no module. What was not attempted is reported as
skipped rather than left to be inferred from silence.
2026-09-28 15:38:16 +02:00
mesh-admin 017b1401e4 Merge pull request 'ADR 0135: a module version prepares its state before it runs' (#161) from decision/0133-0134-migrations-and-deploy-facts into main 2026-09-28 09:57:47 +00:00
jschoubben 476cda417d ADR 0135 supersedes 0133: a module version prepares its state before it runs
Two faults in 0133, both caught on review. It put the declaration on a container — one resource kind
the host applies — so every author would restate the machine's arrangement and a module's own
lifecycle would be tied to how its artifact happens to run. A module declares entrypoints for its
tools and its provisioner; preparing its state is the same vocabulary and nothing about a runtime.

And it derived the scope from the machine, which the facts already answer: a consumer is a module on
a machine (issue 022, migration 0015), so what the mesh provisions is per consumer. A module on three
machines has three databases, there is no shared state to race over, and the lock obligation 0133
invented was for a situation the mesh does not produce. The level question HAL answered with stages
dissolves — the scope of preparation is the scope of the state, and the mesh knows it.

0133 keeps its reasoning and gains a pointer; design 32 and issue 133 name the live record.
2026-09-28 11:57:16 +02:00
mesh-admin 74ba3ff1d4 Merge pull request 'ADR 0133 and 0134: who runs migrations, and the mesh saying what it applied' (#160) from decision/0133-0134-migrations-and-deploy-facts into main 2026-09-28 09:45:49 +00:00
jschoubben 891c7a945e ADR 0133 and 0134: who runs migrations, and the mesh saying what it applied
0133 — a module owns its migrations and the mesh owns when they run. A container declares what must
run before it; the mesh derives the gated step from the resource it precedes, so the image, the
environment and the credentials come from the one place they are described. The module owns the SQL,
the dialect and the lock; the mesh owns the moment and refuses to start a version whose step failed.
Per node, with no level: a step that ran once somewhere leaves every other machine ungated, and
'once, mesh-wide' is what holding a seat already means.

0134 — the pipeline is observable from a merge to an artifact and goes dark at the machine. What a
node now runs, and what it refused, become facts under the control plane's own seat, emitted when
what a machine runs changes rather than on every convergence pass.

Design 32's lifecycle carries both; issue 133 points at them as what ends the matter it opened.
2026-09-28 11:45:47 +02:00
mesh-admin 83cbeb8db8 Merge pull request 'Issue 133: the control plane's schema is migrated at birth and never again' (#159) from issue/133-the-control-plane-migrates-before-it-serves into main 2026-09-28 08:27:32 +00:00
jschoubben ab6db9369b Issue 133: the control plane's schema is migrated at birth and never again
The mesh replaced its own control plane with a build carrying a migration, applied none of it, and
then recorded no build for three quarters of an hour while saying everything was fine. ADR 0052
already prescribes the shape — a run-once step that gates the server — and the control plane was the
one module that did not use it.
2026-09-28 10:27:30 +02:00
mesh-admin 84571b4825 Merge pull request 'ADR 0132: a seat carries the tools its holder must serve' (#158) from decision/0132-a-seat-carries-the-tools-its-holder-must-serve into main 2026-09-28 08:17:01 +00:00
jschoubben d57196102d ADR 0132: a seat carries the tools its holder must serve
A role's tools belong to the role, not to whichever module holds it today: the seat declares them
with their schemas, serving them is a condition of occupying the seat, and what the mesh can do
becomes a read of its own records rather than a question nothing answers. A module keeps its own
tools — the same module may run without the seat, and then only its own name is true.

Design 33 follows: the three families, addressing a node-scoped seat, discovery, and what serves
this to an agent.
2026-09-28 10:16:59 +02:00
mesh-admin e6402cf777 Merge pull request 'Issue 132: a module can be recorded without the directory it lives in' (#157) from issue/132-a-module-can-be-recorded-without-its-directory into main 2026-09-28 07:20:07 +00:00
jschoubben 4f0d144833 Issue 132: a module can be recorded without the directory it lives in
Nine modules could not be rebuilt: their record named the repository and no directory, so every
build looked for a manifest at a repository root that has never had one. Resolved by mesh-controller
— `module add` takes the directory and the forge, and the rule is checked rather than described.
2026-09-28 09:20:05 +02:00
mesh-admin 98d94ef71e Merge pull request 'Design 28: 5.5 done, the mesh has one bus; issue 131 resolved' (#156) from design/28-one-bus-issue-131-resolved into main 2026-09-28 01:59:41 +00:00
jschoubben 4e13280604 Design 28: 5.5 done, the mesh has one bus; issue 131 resolved
The AMQP transport is gone from the control plane and the hosts (mesh-controller #112,
mesh-host #39). On the way: no build had ever recorded its bases, so every bases-first order
walked an empty graph; the builder now reports what it was handed and the graph is read from
builds (mesh-controller #113/#114). Issue 131 is resolved by the forge module's merge event,
the control plane following it, and those edges.
2026-09-28 03:59:39 +02:00
mesh-admin 8783a13448 Merge pull request 'Design 28: the mesh runs on the new bus' (#155) from design/28-the-mesh-runs-on-nats into main 2026-09-28 00:40:29 +00:00
jschoubben a31cfcf461 Design 28: 5.3 is built and was used for the hand-over 2026-09-28 02:28:06 +02:00
jschoubben 75b3861911 Design 28: the mesh runs on the new bus
Tasks 4.3, 5.2 and 5.4 are done as of 2026-09-28 02:25: every machine reports on
the new bus, the seat is held by the module that provides it, the old broker is
unassigned and forgotten, and every credential was minted afresh at the end.

5.2 records what it took, in the order it was found and each fixed on the trunk
before the next step, and how the bootstrap loop was broken once, by hand.
2026-09-28 02:27:49 +02:00
jschoubben 694555214a Merge pull request 'Design 26: which assignment holds a seat is on record, and changes as one act' (#153) from design/26-a-seat-is-held-on-record into main 2026-09-27 21:22:56 +00:00
jschoubben a7249541df Design 26: which assignment holds a seat is on record, and changes as one act
Until now the holder was derived — assigned and claiming — and a second eligible
assignment was refused, so a seat could not pass from one holder to the next
without a moment where nobody held it. The controller finds its own bus through
one of these seats, and that moment took the control plane down on 2026-09-27.

The holder is now a row the controller keeps, written by `seat <name> --to
<node>/<module>` in the same write that removes the previous one. No row means the
old rule, so nothing changes for a mesh that never hands a seat over; with a row,
another eligible assignment is silent rather than refused, which is what lets the
next holder run beside the current one until the switch. A holding is the
assignment's and goes when it does. Each rule names the test that checks it.

Under ADR 0131; design 28 task 5.3 is the work.
2026-09-27 23:20:36 +02:00
jschoubben 95d8253f71 Merge pull request 'ADR 0131: everything on the mesh speaks to the broker seat, and AMQP is not a provision' (#152) from decision/0131-everything-speaks-to-the-broker-seat into main 2026-09-27 21:06:02 +00:00
jschoubben 784b487bf9 ADR 0131: everything on the mesh speaks to the broker seat, and AMQP is not a provision
Taken during the outage of 2026-09-27, when the protocol leaked into the seat's
contract: to hold mesh-broker a module had to provide amqp, so the module that
will carry the bus could not hold the seat that names the bus, while the module
being retired could. Supersedes 0127. Modules depend on the seat and reach the
bus through the sdk; no manifest provides or requires amqp; the old broker's
module and the two modules that required it leave the catalogue; the AMQP
transport is deleted once every node reports on the new bus.

Design 28 step 5 rewritten under it: the seat handover becomes its own task and
is built first, because the seat the control plane dereferences cannot be empty
in between — that emptiness was the outage. The cost note now carries what was
measured rather than what was assumed.

0128 and 0130 extended 0127; each now rests on 0131 with a dated note and
changes nothing it decided. Every other citation of 0127 names its replacement.
records.py still fails on 0120/0112, which predates this branch.
2026-09-27 23:03:05 +02:00
jschoubben 7ae711ba0b Merge pull request 'Building the bus: the decisions the work needed, and what it taught back' (#150) from feat/nats-genesis into main 2026-09-27 17:06:40 +00:00
jschoubben 51739c4302 The order the repositories land in is part of the rollout
Derived while merging and visible from no single repository, so it belongs
written down rather than re-derived later: the client library before the
catalogue, because it is where the subject is derived and a converted module
against the old one publishes the local name itself; the catalogue before the
controller, because the controller refuses an old-style name outright and would
make every unconverted module unregisterable.

Two tested properties are what make it safe and not merely ordered. An old-style
name passes through the derivation untouched, so an unconverted module keeps
working at every step. And a converted name derives to exactly the key the old bus
published, so nothing moves on the wire until 5.2 sets the variable.

The failure avoided is 127's own, which is why this is worth a table: a publisher
and a subscriber disagreeing about a subject log nothing anywhere.
2026-09-27 19:04:45 +02:00
jschoubben 5cac268457 Two checks were wrong about where a seat is judged
Correcting design 26 to match what the merged code does, and fixing issue 112's
status, which used a word the vocabulary does not have.

A claim on a seat the manifest does not itself declare is refused at registration
rather than by the parser. A module may hold a seat another module declared —
that is why ADR 0126 has a caller name the seat and not its provider — so whether
the name exists is a fact about the whole catalogue.

And a declared seat may promise nothing. That is a marker seat, and most
node-scoped seats are markers: which module is this machine's packet filter. ADR
0126's "a declared seat carries a protocol" says what a holder must satisfy, not
that every seat offers something.

`records.py` still fails on ADR 0120 resting on a proposed ADR 0112, which is not
this branch's and not mine to decide.
2026-09-27 18:57:49 +02:00
jschoubben ce6ae943b7 Merge main: renumber this branch's records around the trunk's
Both lines of work numbered from the same point, so four decision records and one design
document existed twice with different content. The trunk keeps its numbers and this branch
yields — the only rule that scales, because the trunk's are already cited by what merged
before them.

  0117 the bus is the only broker        -> 0125
  0118 a module declares its own seats   -> 0126
  0119 amqp is a provision, not the bus  -> 0127
  0120 the mesh bus is required          -> 0128
  0123 a seat carries its role's protocol -> 0129
  0124 the predecessor is ending          -> 0130
  design 29, what a module declares       -> design 32

Applied to the code repositories too, because a stale reference is worse when numbers
collide than when they dangle: the reader lands on a real record that decided something
else.

Two reconciliations the merge forced, both real:

**0110 was marked wholly superseded and was not.** Its successor says in as many words that
everything 0110 decided about what a seat *is* stands untouched — and two records that
landed on the trunk rest on exactly that part. So it is accepted again, extended rather than
replaced, with a note saying which of its claims moved and where.

**A seat's protocol becomes columns, not fields.** The trunk moved the seat set out of
compiled code into a table the controller owns. This branch had added what a role accepts,
emits and serves to the Go slice. The decision is unaffected and the mechanism is better for
it: giving a role a protocol is now a write rather than a rebuild, which is the trunk's own
argument applied to what this branch added.

One check still fails and it fails on main too: a record resting on ADR 0112 while that is
still 'proposed'. Left alone — it is not this merge's to answer.
2026-09-27 18:23:41 +02:00
jschoubben 9ef9830dcf ADR 0122: the predecessor is ending, and its broker goes with it
ADR 0119 rejected giving the old broker a retirement condition and said why: "its
clients are not only the predecessor's, so the retirement condition describes a day
that will not come". The operator has said that day is coming — the predecessor is
deprecated, some of it still running, none of it being migrated, left to stop rather
than moved.

Recorded because three documents reason from the premise it overturns. Design 25 §5's
"no day anything is waiting for", §9's "the predecessor's clients never notice", and
design 28's closing note that the predecessor's world does not need to move.

**And it needs no new machinery, which is 0119 being paid off rather than revised.**
Because that record made the broker an ordinary provider rather than a compatibility
module, ending it is unassigning a provider whose provision nothing requires — something
the module system has done since it existed. So step 5.3 finishes instead of trailing
off, and the transitional double announcement of a build outcome has a date.

The consequence worth planning around: the predecessor's own mesh talks over that
broker, so shutting it down ends the tooling that reaches this installation's machines
from a workstation. The rollout is driven from the node, or driven before the broker
stops. That is a sequencing constraint on 5.2, not an afterthought.

What survives is `amqp` as a provision: a module that genuinely needs an AMQP broker can
still be given one. What retires is this broker's role as the predecessor's.
2026-09-27 18:11:11 +02:00
jschoubben dfac01a6dd 5.2: the readiness half is in, and what a failed move actually costs
`rollout check` answers from records whether this mesh could move its bus, and names the
next step for each thing missing. The move itself waits on that check having been run
against a real mesh — writing the irreversible half before its question has ever been
asked of something real breaks the plan's own rule about beds by another route.

And the cost of being wrong is written down rather than assumed: the old broker stays for
its other clients, nothing in a served request's path goes over the mesh's own bus, and
what a failed move costs is the mesh's ability to change things rather than the services
its modules serve.
2026-09-27 17:59:26 +02:00
jschoubben 51f5ce5c0f 4.5 done: the catch-up needed nothing built, which was the answer
Three ways to do the replay were weighed and the right answer was that the bus being
moved to already does it. A queue on the old bus receives only what is published after
it is bound, so everything built before the catalogue existed was announced to nobody. A
stream is a log and a consumer is a position in it: a consumer created later starts at
the beginning and the builds are simply there. Checked against a running server, because
the decision rested on it.

So "who replays" has no answer because nothing replays. The mechanism was never about
builds — it was about a queue that could not remember, and carrying it across would have
carried a workaround for a limitation that no longer exists, with nothing looking wrong.

Retiring it belongs to step 5, with the rest of what only the old bus needs.
2026-09-27 17:23:49 +02:00
jschoubben fc64a2c1a4 4.4 done: a person's account and their client
The account existed as a permission model and as nothing a person could be given; there
is a record and three commands now. The client is two surfaces over one thing, a command
line and an MCP server, both using the client a module's runtime uses — so what a person
may do is answered by the same permission list that answers it for a module.

Design 25 §7 says nothing of this is built before its bed passes, and this was built
before. Noted in the task rather than quietly ignored.
2026-09-27 17:07:28 +02:00
jschoubben 9fc4e74b0d Merge pull request 'to-be 31: a module declares its fail2ban jail, mesh composes them per node' (#149) from design/a-module-declares-its-jail into main 2026-09-27 15:00:16 +00:00
jschoubben 225dfa9451 to-be 31: a module declares its fail2ban jail, mesh composes them per node
A node's intrusion filter should be composed from its assigned modules, like
its firewall (the Filtering mechanism): a service module (postgres, mssql,
mailu) declares its jail in its manifest (filter + stanza, no node/path per ADR
0112), and the mesh writes the jails of a node's modules into the fail2ban
holder's jail.d. The base (sshd, recidive, ignoreip=mesh-range) stays the
fail2ban module's. Records the model after novox's HAL per-module jails were
lost as dangling symlinks; the ignoreip is now safe on disk, the service jails
need this to be restored.
2026-09-27 16:59:56 +02:00
jschoubben bfefb1dbdb 4.3: the installer can raise a mesh on the new bus
A foundation template that stands the server up, writes its settings and the mesh's
first user list beside them, and starts a controller on the new bus. The first user list
is the installer's because at genesis there is no mesh to compose one — a bootstrap
credential, rotated like the store's.

The carried list is checked against what the controller derives, since a mesh cannot be
raised twice to discover they disagreed. That check immediately found the composer
granting a role's whole event branch as well as the one event it follows.

What is left of 4.3 is running it, which is 4.1's bed.
2026-09-27 16:39:48 +02:00
jschoubben 85c0a3e567 4.2 done: a build is work submitted to a role, on both buses
Both sides behind a seam, one implementation per bus, and the outcome is the role's
own event so one publish reaches the asker, the controller and the catalogue. Checked
against a running server, including the part the decision rests on: a third party
hears the same outcome the asker does.
2026-09-27 16:02:17 +02:00
jschoubben a39765c924 Merge pull request 'ADR 0122: a seat is data the controller owns; a rename is a database update' (#148) from design/a-seat-is-data into main 2026-09-27 13:41:10 +00:00
jschoubben a8921fe737 ADR 0122: a seat is data the controller owns; a rename is a database update
Reviews 0110/0121 after a session where renaming seats cost three freezes, a
builder deadlock, and hand-resolved manifests. The seat rules were right; the
set being a compiled Go slice referenced by name-string everywhere was the
mistake. Seats become a table keyed by a stable id; claims/held/production code
reference the id; a rename is one UPDATE, no rebuild, no re-registration, no
freeze. The build machine reads the set from the mesh instead of embedding it,
removing the controller/builder seat coupling. Closed set and scope naming
unchanged; only storage and reference change. Outstanding renames (registry
seats, private-network scope) wait for this — as data each is a write.
2026-09-27 15:40:55 +02:00
jschoubben c8f430290e ADR 0121: a seat carries the protocol of its role
The mesh's own seats said who does a job and nothing about what may be said to
them or by them, and that gap showed up three times in one day looking like three
different problems: a build machine with three audiences for one outcome and no way
to derive a grant for any of them; an event genuinely about a role with nowhere to
live but the namespace of whichever module holds that role today; and a catalogue
catching up on builds, where every option needed a grant the design refuses.

One cause — the mesh has roles it cannot describe. So the `mesh-*` seats take the
same three fields a module's seat has, and the machinery that already derives
authority, queues and consumers from a declared seat does it for these too.

Builds become work submitted to a role, and `mesh.build.request`,
`mesh.control.built` and the BUILDS stream retire. A work queue shared by several
build machines is exactly what a seat's `accepts` is, so a second mechanism for it
was two places a permission could be wrong. The outcome is the seat's own event,
which means one publish still reaches whoever asked, the controller that records it
and the catalogue that places it — the fan-out a shared exchange gave for free,
written as a subject the mesh derived rather than a topology somebody configured.

That also avoids the grant that ruled out the alternatives: no holder needs
permission to publish into an asker's inbox.

The blocking gap is now named rather than incidental: the shared library has no way
for a module to publish on a seat. The build machine is Go and reaches the bus
directly, so it is unaffected; the artifact-store event waits.
2026-09-27 15:24:17 +02:00
jschoubben d98d6fca11 Merge pull request 'to-be 30: the mesh updates itself on a push' (#147) from design/the-mesh-updates-itself into main 2026-09-27 12:49:32 +00:00
jschoubben b8cdfce16d to-be 30: the mesh updates itself on a push
Records the manual update process (module moved -> build -> reconcile; and the
breaking-change freeze/re-register recovery), and the two things that make
self-update more than a webhook: the build-on-push trigger is currently HAL's
(hal-gitea-tools on :9877), a retirement gap the mesh must replace with its own
forge-webhook trigger wired to every repo including mesh-controller; and the
builder validates manifests too, so a breaking change couples controller +
builder + manifests + hosts, and renaming the builder's own seat deadlocks its
rebuild. Names the transition discipline (accept old+new for one release) that
self-update needs so a push does not auto-freeze.
2026-09-27 14:49:02 +02:00
jschoubben c4a8455e2e Issue 127 resolved; design 29 says what wildcards are and how the rule is checked
Every module named its events the way the old bus spelled a routing key, so on the
new bus every cross-module subscription pointed at a namespace nobody publishes to.
Nothing failed — the services started and none of them reacted. Converted, and the
rule now has checks at both scales: at registration for one manifest, and as a test
across the whole catalogue where a consumed event's emitter is present.

It was larger than the report said, in two directions nobody had looked. Forty-three
files of module code pass the event name at runtime, so the code mattered as much as
the manifests. And both clients had to learn the mapping — without that, converting
the modules would have broken the mesh that is actually running, which is the
opposite of what fixing this was for.

Design 29 gained three things it did not say: what a wildcard is (`*` for one name,
`**` for the rest, spelled the mesh's way and derived to each bus's own), that an
event about a role belongs on the seat and why that is not yet possible, and how the
rule is checked — because "a subscription that matches nothing is silence" is exactly
why nobody noticed thirty-seven manifests being wrong the same way.

4.2 and 4.3 are unblocked. The catch-up half of 4.5 is not: it is a decision, and it
narrowed rather than closed. It cannot be a reply to a module's inbox, because that
needs the blanket grant design 25 §4 refuses.
2026-09-27 14:45:08 +02:00
jschoubben bf7297b7ed Merge pull request 'ADR 0121: keep distribution, retire only verdaccio; node-* seats migrated' (#146) from design/0121-keep-distribution-fix into main 2026-09-27 12:36:54 +00:00
jschoubben 1bb0ef5658 ADR 0121: keep distribution, retire only verdaccio; node-* seats migrated
Records the reversal: distribution stays as the mesh's OCI registry (it serves
every artifact-store:// image); only verdaccio, a redundant second npm registry,
is removed. The 'consolidate onto gitea / retire distribution' direction was
dropped. Also records that the node-* rename was executed as one controlled
migration with a brief compose freeze, and why the delivering registry seats
are deferred rather than folded in.
2026-09-27 14:36:33 +02:00
jschoubben 5a9917f802 Merge pull request 'ADR 0121: a system seat is named for its scope; a module may define its own' (#145) from design/system-seats-are-named-by-scope into main 2026-09-27 12:15:22 +00:00
jschoubben 8a6ee9177c ADR 0121: a system seat is named for its scope; a module may define its own
The control plane's seats grew a second naming style (the-*) beside mesh-*,
and the closed set was the only place any seat could be defined. This settles
both: system seats are mesh-* (one, mesh-wide) or node-* (one per node), named
for scope; a module may define its own seat outside the closed set. Folds in
the seat review: mesh-build-machine (scope fix), mesh-private-network (one
server + client modules, dropping per-node VPN choice), showcase becomes the
first module-defined seat, node-uplink, and the node-* renames — plus the
registry consolidation onto gitea, which reshapes the registry seats and gates
retiring distribution/verdaccio. Records why the renames are a coordinated
migration and why distribution cannot be removed until gitea serves images.
2026-09-27 14:15:05 +02:00
jschoubben dd577e9ebe 1.7 is done; 4.1 waits on nothing but the bed
Minting on both halves, the file delivered per push, and the bus's objects
asserted on every start — verified against a real server that asserting twice
changes nothing, that a machine joining an already-raised bus is accepted, that
each node's consumer is bound to its own declaration subject, and that CONTROL
does not dead-letter before the controller gives up.

Which bus the mesh is on is one fact, and being told about both is refused at
start rather than warned about: a mesh half on each is one where a declaration
goes out on one bus and the report comes back on the other while every component
logs success — ADR 0074's failure arriving through configuration rather than code.

So 4.1 no longer waits on code. Both links speak NATS, the composition happens,
and every claim behind them has a unit test or a check against a running server.
What none of those can stand in for is a mesh raising itself, which is what the bed
is — this is where the code stops and the lab starts.
2026-09-27 03:20:42 +02:00
jschoubben a9b8f41570 Design 25 §4: the server verifies no client certificate, and writes only accounts
Two corrections of fact, both found by building the module's image and connecting
to it as a host would.

The first composed configuration said `verify: true`, which makes the server
demand a *client* certificate — and nothing in the mesh presents one. A host pins
this server's exact certificate and authenticates with the password the mesh
minted, and so does a module's runtime. Every connection in the mesh would have
died at the TLS handshake before any password was looked at, with an error that
reads as a fault in the client. TLS is still required; verify only decides whether
client certificates are checked. Mutual TLS is a later question and would need
machinery the mesh does not have — a certificate per module per node.

And §4 read as though the controller wrote the whole file. It writes the user list
and nothing else: ports, TLS paths and a store directory belong to the container
the module raises. The two files share one directory of necessity, because an
absolute include path is resolved relative to the including file's own directory.

The decision stands in both cases — accounts are composed, not called for, and
passwords are minted and sealed. What changed is what the file says and who writes
which half.
2026-09-27 02:51:40 +02:00
jschoubben 92a5e8fc05 Merge pull request 'ADR 0120: a roster fact carries its format as a template; rewrite to-be 29' (#144) from design/roster-fact-is-a-template into main 2026-09-26 23:51:21 +00:00
jschoubben 8c91ba1cfa 1.7 in progress: the list is derived and the keys are kept
What is in: a bus user's hash is recorded and its plaintext returned once, and
the user list is derived from the machines, what each runs, every manifest and
which machines hold a live token. Permissions stay derived rather than stored,
because a stored copy could disagree with the records it came from while both
looked internally consistent.

What is out, with what each needs, so the next person does not rediscover it:
delivery, which has one open question about what a module declares in order to
receive the file — design 29's ground, not this document's; minting, which is
transport-coupled because an enrolment reply carries one password and a node on
the old bus must not be handed a credential for the new one; and calling the
assertions from a start path.
2026-09-27 01:50:04 +02:00
jschoubben 0f7f628730 ADR 0120: note the shared/region interaction with hq 128
A roster fact may be shared — written into a marked region of the machine's
file (into: block, hq 128) rather than as the whole file. The template
renders the content; shared decides how the host lays it down. Composes with
hq 128: the region mechanism is the host's, the format is the module's.
2026-09-27 01:42:00 +02:00
jschoubben 970da74136 1.7: the composer exists and the composition does not
Correcting a tick and a claim I made one commit ago. 4.1 does not wait on an
enrolment user per live token; it waits on the whole composition, of which that
user is one input.

Tasks 1.3 and 1.4 are honest about what they built — the composer, the
derivation, the permission model, the stream and consumer definitions, the
asserter, all pure and held by unit tests and a golden composition. Nobody wrote
the caller. Measured: outside the package that defines them there is not one use
of the composer, the permission derivation, the stream set, the stream asserter or
the principal type. Step 1's "done when" claims every account and permission
composed from the manifests, and a mesh raised today would stand up a server with
no user list at all.

It also needs state the mesh does not keep. Design 25 §4 says the file holds
bcrypt hashes and that passwords are minted and sealed exactly as today — but
today the mesh mints one, hands it to the broker through a management call, seals
the plaintext to the holder and keeps nothing. With no management call the hash
has to survive every later recomposition, because the first thing a new module or
a person's access change touches is a file that must still hold every other
user's password. No bcrypt hash is stored anywhere in the controller.

Named as its own task rather than folded into 1.3, so the gap between "the parts
of step 1 exist" and "the mesh does any of it" is visible.
2026-09-27 01:39:20 +02:00
jschoubben 4d4012cdf6 ADR 0120: a roster fact carries its format as a template; rewrite to-be 29 around it
The facts mechanism formatted the roster in Go in the control plane — one
formatter per fact, in the consumer's own configuration language. ADR 0120
makes a fact a path and a template: the mesh owns the data, the module owns
the format, and the control plane holds no format at all.

to-be 29 (operator accounts + what lives under a home) is rewritten to ride
it: the ssh files become roster templates, the whole ~/.ssh is owned with a
found/owned boundary that cannot lock the operator out, keys are mesh-owned
through an SSH CA (existing keys adopted not regenerated, the operator's
personal key signed not minted), and the ssh-agent is a user-scoped service.
2026-09-27 01:33:12 +02:00
jschoubben 0a61531c42 WBS 3.5 done; 4.1 waits on one thing, an enrolment user per token
All three halves of the host's link are through seams, and the reply address in
the payload is now proved from both ends rather than one — the test asserts the
transport's own field held the consumer's ack address by the time the request
arrived, so a server that stopped claiming it fails a test instead of letting the
reason become folklore.

Two things had to be built for the host to hear anything at all: a node's
declaration consumer, which only the controller may create, and the enrolment
user's inbox, which design 25 §6 names and the composer granted none of. Both
were silent gaps — a node with either missing looks correct and hears nothing.

What remains is a single piece: something that composes an enrolment user per
live token. On the old bus that account is made imperatively through the broker's
management API; here there is no management API, so issuing a token has to
recompose the server's configuration. It is the only thing between the two links
and a mesh raised on NATS from nothing, so 4.1 now says so.
2026-09-27 01:32:26 +02:00
jschoubben 4b0b659084 Merge pull request 'ADR 0119: a taken tunnel's predecessor is retired once the take is proven' (#143) from decision/0119-a-taken-tunnels-predecessor-is-retired into main 2026-09-26 22:59:15 +00:00
jschoubben f555d523c7 WBS 3.4 is done both halves; issue 127 holds 4.2, 4.3 and catch-up
The controller's inbound is through a seam with both transports behind it, and
the store window is now the server's rather than the controller's memory. Seven
claims about that were asked of a running server rather than reasoned.

Wiring the controller's own subscription is what found issue 127: every event
name in the catalogue is still written the way a routing key is, so design 29's
derivation turns a consumer's declaration into a subject no emitter publishes.
Thirty-seven manifests, one that cannot be composed at all. It fails on the first
mesh raised on the new bus and not before, which is why nothing had caught it —
the conformance fixtures pin one emitter against one subject, and both halves of
that pair are correct.

The node-facing flows are unaffected: those subjects are the mesh's own and
derive from nothing a module declares.
2026-09-27 00:55:07 +02:00
jschoubben 4f93d304d7 WBS: a person's account is done; the client is not blocked by step 3 2026-09-27 00:18:05 +02:00
jschoubben d898bd87e8 Step 5.4 was wrong from 0119 onward; removed
It waited on a retirement condition 0119 abolished when it made the
deprecated broker an ordinary provider. A step waiting for a condition
nobody set would sit open forever.
2026-09-27 00:16:45 +02:00
jschoubben 7e4da874a9 Design 25 §2: the eaten reply address is verified, not assumed
A claim the whole enrolment handshake rests on, now measured against a
running server rather than reasoned from documentation — and held by a test
so it cannot become folklore if a server version changes.
2026-09-27 00:15:00 +02:00
jschoubben 53092020eb WBS: asking a tool is through the seam; a build is a different shape 2026-09-27 00:11:44 +02:00
jschoubben 00817fb9e3 Design 25: the store window, and what moving it into the server changes
The guarantee is the same and the mechanism is simpler — a nak with a
delay, no parked list, nothing lost when the controller restarts. It costs
one thing: a naked message comes back whatever happened meanwhile, so an
older report is redelivered after a newer was applied. A report already
carries the digest of the declaration it answers, so supersession becomes a
check rather than memory — ordering settled by what a message says, not by
when it arrived.
2026-09-27 00:06:15 +02:00
jschoubben 9946e852e1 WBS: 3.5's outbound half is in 2026-09-27 00:01:56 +02:00
jschoubben 9f6aa7ea9c The bus is the mesh's centre, not a transport that replaced one
Two things. A paragraph from the superseded 0117 survived beside the 0119
correction that reversed it, so §5 said both that the amqp interface
retires and that it does not. The stale one is gone.

And the framing. §1 opened with "the bus carries five kinds of traffic
today, and this design keeps the five", with a column mapping each to the
queue it used to be — which describes the mesh's nervous system as a port
of something that did a fraction of this. It now says what the bus is: a
role addressable without knowing its holder, the mesh's own state, work
that queues until somebody can do it, and permissions derived from what a
module declared. Conditions, observation and a person's client land there
too as they are built.

Glossary gains `bus` and `the deprecated broker`, with a note on why not to
say "compatibility broker" or name it after a protocol — the second invites
exactly the backwards framing this commit removes.
2026-09-26 23:54:09 +02:00
jschoubben 672c994afa WBS: 3.4's seam is in, outbound half through it 2026-09-26 23:47:40 +02:00
jschoubben 2f9bb73685 WBS: the first fixtures are in, and what byte-for-byte means 2026-09-26 23:41:15 +02:00
jschoubben 5c193b3f54 WBS: 3.8's check is written 2026-09-26 23:34:42 +02:00
jschoubben fbf9440e8e WBS: 3.7 and 3.8 done, with the one check 3.8 still owes 2026-09-26 23:34:04 +02:00
jschoubben 9510bf5311 WBS: 3.2 done 2026-09-26 23:33:10 +02:00
jschoubben d940e14ec8 Design 19: the protocol on NATS
Task 3.2. ADR 0074's model is untouched — floor plus capabilities, partial
implementations legitimate, identity from the credential, dedup on
x-event-id, conformance as executable fixtures. The transport beneath it is
rewritten: exchanges and queues become subjects and streams.

Statements marked *verified* were checked against a running server while
the runtime's client was written, not reasoned from documentation. Three
of them are things the specification would otherwise have got wrong:

- the payload is the body alone, with metadata in NATS headers; an
  implementation that nested the whole envelope would agree with nobody
- a durable name may not contain a dot, while the ack subject joins two
  names with one — conflating them looks right in a permission list and is
  refused as a consumer name
- a certificate must carry a name the bus is dialled by, because the NATS
  client has no hook to replace hostname verification the way pinning did
  on AMQP

And one limitation lifts: a module may now call another's tool. Issue 049
recorded that a scoped account could not declare the reply queue a caller
needs, and ADR 0095 routed every ask through the control plane because of
it. Per-account inbox prefixes plus allow_responses replace that. ADR 0095
is not reversed — the control plane is still how a person asks — but
module-to-module calling stops being a question about capability and
becomes one about policy, which `uses` already answers.
2026-09-26 23:32:48 +02:00
jschoubben 80456981be WBS: 3.6 done, and the certificate constraint it surfaced
The NATS client has no checkServerIdentity hook, so pinning no longer makes
the name check redundant — the bus's certificate must carry a SAN matching
the address nodes dial.
2026-09-26 23:29:56 +02:00
jschoubben d2ed3152d3 Seat renames done; 0118 was wrong that it was a migration
A holding is derived at resolution from manifests, never stored, so there
are no recorded old names to rewrite. The work is an edit plus a kept rename
table — kept because a module lives in its own repository and may be
registered long after the catalogue stopped using an old name.
2026-09-26 23:06:55 +02:00
jschoubben 24d99ddd24 WBS: 3.9 done, and 1.4's client with it 2026-09-26 22:28:59 +02:00
jschoubben b759e36bfd Design 29: tools, not serves; WBS 3.9 partly done
The manifest already uses serves for a provision's facts, so a module's
tools take their own key. Declaring them is itself new — until now a
module's tools existed only in a runtime environment variable.
2026-09-26 22:17:17 +02:00
jschoubben 0b8e84334f Why module events share one stream, checked against the server
Storage is not a property of a subject — a stream is a separate object that
covers one — so the question is always how many streams, not which topics
are durable.

Three facts decide it, two of them verified rather than assumed: NATS
refuses overlapping streams instead of merging them, so a shared stream
plus a per-module one is not available at all; a filter cannot express an
exception; and a stream per module turns one cross-module consumer into one
per module. So one stream, with per-subject caps for the fairness that
matters. Per-module age is genuinely unavailable, and a module that needs
it declares a seat.
2026-09-26 21:49:55 +02:00
jschoubben 39c802cbd4 Step 2: adoption recreates the bus once, on purpose
2.3 was already true and is now proved — the seat refusal is generic, and
three tests pin what matters: a second bus is refused by name, a different
bus implementation is refused too (which is what makes the bus replaceable),
and the AMQP broker no longer contends so both run on one mesh.

2.1/2.2 turned out not to be a no-op. The host keeps a container only when
its spec matches exactly; genesis raises the upstream image and the module
declares the mesh-built one carrying the entrypoint, so assigning it
recreates the container. That is ADR 0067's pivot and it is safe only
because the bus carries nothing yet — which is why step 2 comes before
anything speaks NATS. After it, never again: the config is a directory
mount, so accounts change without touching the container's spec.
2026-09-26 21:40:02 +02:00
jschoubben 86a084b7ff The mesh bus is required, not ambient
Design 29 said no module requires the bus. The catalogue disagrees: 49 of
72 modules take a broker credential and 23 do not, so an ambient connection
mints an account for a third of the catalogue that never speaks — and the
49 each hand-write the path it lands at, which is provisioning done badly
by hand.

The bootstrap argument that made it ambient was narrower than it looked.
"A provisioner needs an account before it can run" is true of a provisioner
process and says nothing about a provision the controller answers, and the
controller is not waiting on a bus account to compose one.

So: the mesh-broker seat delivers mesh-bus; a module requires it and gets an
address, a sealed credential and the trust to verify the server; a module
that requires nothing has no account at all. The requirement delivers the
connection, the declarations shape the authority, and declaring a subject
without requiring the bus is refused as incoherent.

mesh-bus and nats are deliberately two names: a module may run its own NATS
as a backing service exactly as one provides amqp, and a manifest saying
"nats" would otherwise mean either the mesh's nervous system or a private
queue.

The seat's Delivers was wrong twice today — amqp, then empty — and the
comment says so rather than reading as though it were always right.
2026-09-26 21:17:21 +02:00
jschoubben 85f972749a AMQP is a provision, not the bus
0117 went a step further than it had grounds for. It was right that the bus
is the only bus, and wrong that the amqp interface must therefore retire —
because it conflated two reasons to want a broker. Using one to reach
another module is a second bus and stays refused. Needing an AMQP broker as
a backing service, the way something needs a database, is ordinary, and
forbidding it would make the mesh unable to run normal software while
calling that architecture.

So the broker becomes a plain provider module: no seat, not foundation,
never raised at genesis, no retirement condition. lavinmq now claims nothing
and provides amqp; nats claims mesh-broker and provides nothing.

The rule that survives is about direction, not software: inter-module
communication goes over the bus. A module may hold a broker for itself; it
may not use one as a channel to another module. That is a review judgement
where 0117 could have used a parser, which is the honest cost.

0106's progressive insight was itself wrong and is corrected by a second one
there — nothing moves off the old broker, so its "one purpose" sentence does
not become true, it is just not what that server is.

The insight check needed two fixes it found itself: a date may carry
trailing words, and a bold run with a link is discussing an insight rather
than marking one. All four bad shapes still fire.
2026-09-26 21:08:40 +02:00
jschoubben 7abb268de6 Step 1 done but for its bed
1.1 to 1.4 built and tested. 1.5 turned out to need no controller change:
it already resolves the broker by seat and names no broker module in its
source, which is what ADR 0079 was for. The genesis module set naming is
scenario and installer config, carried with the bed.

Recorded what must NOT change yet: the amqps:// credential shape and the
5671 default are correct until the rollout, because steps 1-4 leave every
node on AMQP.
2026-09-26 21:03:35 +02:00
jschoubben 78a2274baf Designs 25 and 29 disagreed about the subject space; implementing found it
29 put a module's events and tools in one namespace, 25 kept mesh.events.*
and mesh.tools.*. One namespace is right — a module's authority over its own
name becomes a single pattern the server enforces — but it needs a kind
token, because a stream is a subject filter and mesh.mod.*.> would persist
every tool call in the mesh. Tools stay on core NATS for the reason 25
already gives.

So: mesh.mod.<module>.event.<name>, .tool.<name>, and seats the same shape.
2026-09-26 21:02:18 +02:00
jschoubben 3f9b316015 Design 25: a scoped inbox needs allow_responses, or nothing can answer
Found composing the first real configuration. Scoping every inbox to its
owner is right and leaves a responder unable to reply, because the answer
goes to the caller's inbox. The fix is not a wider grant but the server's
own allow_responses: one reply to the subject of a message the user actually
received. Without it every tool call times out while the permission list
looks correct.
2026-09-26 20:58:34 +02:00
jschoubben e05825a881 Design 29: versioning, provisioning and secrets on the bus
Versioning: additive is free; a breaking change is refused while callers are
bound, and the refusal names them, because the mesh already holds the uses
graph; a real break versions the subject, not the seat name, so the role
does not fork; binding is a recorded pin, not a drift to whatever is newest.
Semantic change stays open — no fingerprint sees it, and saying so beats
implying the check is complete.

Provisioning: a provisioner's create/remove/holds IS a serves protocol, so a
provision interface is a seat that also delivers a credential — which is why
design 26 already allowed that. The per-consumer resource is what stops the
two collapsing into one.

Secrets: sealed, so the bus is never trusted with plaintext — but sealed is
not enough, because a stream persists and a durable ciphertext is an archive
the day a key leaks. So a secret never enters a stream: core request/reply
only, and a declaration names a secret rather than carrying one, which is
0098's fetch-don't-store applied where carrying is worst. The vault's own
credential and the bus's own accounts are the two bootstrap exceptions,
resolved the way 0067 resolves the control plane.

Also rewrote the addresses paragraph, which was too compressed to follow:
on-bus addresses disappear because nothing stores them, off-bus ones are
untouched and still 0098's problem, and the bus's own address is the one
that cannot be a subject.
2026-09-26 20:44:16 +02:00
jschoubben 7b4916e9ec Modules declare their own seats; the mesh reserves mesh-*
The architecture 0117 opened needs a module to offer a service as a role on
the bus — one holder, addressed by what it does. A closed table in the
controller cannot express that: a capability a module contributes would
require changing the mesh itself.

But 0110 closed the set for a good reason — nothing could say what seats a
mesh had, and the hand count came out at eleven of thirteen. That argues for
enumerable, not hardcoded, and 0110 weighed free-form against a fixed table
without considering a third option: closed at any moment and derived from
the catalogue. A derived list cannot drift, which is how the count broke.

So: the mesh's seats stay the mesh's, reserved by the mesh- prefix so the
prefix is the rule and there is no list to maintain; ten seats are renamed
to restore 0079's convention; everything 0110 decided about what a seat IS
survives untouched.

Design 29 carries the declaration model: three namespaces, subjects derived
from local names so a manifest survives the wire changing, queues never
declared, five relationships (the job and state shapes 0041 had no room
for), and the build-publish-deploy lifecycle with hard, soft and build-time
dependencies distinguished.

0041 gets a progressive insight: "no per-consumer setup, only a
subscription" was a fact about a topic exchange, and a JetStream durable
consumer is a real object someone creates.

WBS 1.3/1.4 were wrong and say so: streams come at registration and
consumers at assignment, so only the foundation set belongs at genesis.
2026-09-26 20:34:32 +02:00
jschoubben b5b68e8852 Design 25: a host directory bind, not a named volume
Issue 115 is resolved and converted four modules away from named volumes;
the bus's own data is not the place to bring one back. Also: NATS carries
TLS on the client port rather than beside a plaintext one, so there is no
5671/5672 pair to mirror.
2026-09-26 19:34:54 +02:00
jschoubben 814c9e563f The bus is the only broker; step 1 starts
A module does declare requirements the provisioner fulfils — but the broker
it gets that way is a private vhost, the analog of a database, not the
mesh's bus. Two modules of the new mesh depend on it, so the compatibility
broker was never single-purpose and its retirement would have stranded them.

NATS is the heart: one bus, a module's messaging is subjects on it scoped by
what it declares, and no module is handed a server of its own. The seat
delivers nothing; the interface retires with the broker. Also closes the
EVENTS question — one stream, on the bootstrap argument, not preference.

Designs 25 and 28 go in-progress: step 1 is starting.

The insight check caught a false positive on its own first real use — its
bold-run pattern crossed newlines and joined an unrelated `**` to the
marker. Constrained to one line, still catching all four bad shapes.
2026-09-26 19:26:07 +02:00
133 changed files with 10515 additions and 276 deletions
+3 -1
View File
@@ -54,4 +54,6 @@ whose failure has never been observed is a guess about its own correctness.
The development cycle, checked ([ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md)): The development cycle, checked ([ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md)):
a to-be design names a decision, an in-progress/implemented design names its owning code, a a to-be design names a decision, an in-progress/implemented design names its owning code, a
located/fixed issue names its owner, a fixed/resolved issue says what fixed it, a graduated located/fixed issue names its owner, a fixed/resolved issue says what fixed it, a graduated
research overview says what it became. `python3 00-META/checks/cycle.py` research overview says what it became, and no two issue records share a number (issue 155 — the
number is how a record is cited, and `main` lags every open pull request, so two people reading it
allocate the same one). `python3 00-META/checks/cycle.py`
+23 -1
View File
@@ -14,7 +14,8 @@ What is enforced:
its owning code (`code:`) -- no development without a design that says where. its owning code (`code:`) -- no development without a design that says where.
issues a known `status:`; once `located`, `located-in:` names the owner; issues a known `status:`; once `located`, `located-in:` names the owner;
once `resolved`, `fixed-by:` says what fixed it (prose counts -- once `resolved`, `fixed-by:` says what fixed it (prose counts --
"nothing, the capability existed" is an answer). "nothing, the capability existed" is an answer). And no two records share a
number -- the number is how a record is cited.
research a known `status:`; a `graduated` overview says what it `became:`, and every research a known `status:`; a `graduated` overview says what it `became:`, and every
target it names exists. target it names exists.
decisions every accepted record is REACHABLE from the cycle: cited by a design doc's decisions every accepted record is REACHABLE from the cycle: cited by a design doc's
@@ -109,6 +110,27 @@ def main():
"without a design that says where" % status) "without a design that says where" % status)
# ---- issues ------------------------------------------------------------------------ # ---- issues ------------------------------------------------------------------------
# Two records may not share a number. Numbers are taken as "next free after main", and work
# sits on unmerged branches for days -- so two people reading the same main allocate the same
# number, and nothing said so. It happened twice in one evening between two machines, and the
# second collision landed on main with all three checks passing (issue 155). An issue number is
# how every other record cites this one; two records answering to it means a pointer that
# resolves to whichever the reader happened to open.
seen = {}
for folder in sorted(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", ""))):
name = os.path.basename(os.path.normpath(folder))
number = name.split("-", 1)[0]
if not number.isdigit():
continue
if number in seen:
bad(os.path.join("04-ISSUES", name),
"is numbered %s, and so is %s -- an issue number is how it is cited, and two "
"records answering to one means a citation that resolves to whichever the reader "
"opened. Take the next free number across main AND every open pull request"
% (number, seen[number]))
else:
seen[number] = name
for path in sorted(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", "00-report.md"))): for path in sorted(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", "00-report.md"))):
front = frontmatter(path) front = frontmatter(path)
if front is None: if front is None:
+36 -6
View File
@@ -164,6 +164,12 @@ def check_rests_on(failures, records):
# decision is exactly what as-is is for." # decision is exactly what as-is is for."
if rel(path).startswith("03-DESIGN/00-as-is/"): if rel(path).startswith("03-DESIGN/00-as-is/"):
continue continue
# A withdrawn record's citations are history. It instructs nobody -- every reader
# is sent to its superseder -- so what it was built on may itself be withdrawn.
# Refusing that would mean rewriting the lineage of a record whose reasoning is
# the thing the immutability rule protects.
if frontmatter(read(path)).get("status") == "superseded":
continue
# An extension that supersedes legitimately names what it replaced. # An extension that supersedes legitimately names what it replaced.
this = ADR_FILE.match(os.path.basename(path)) this = ADR_FILE.match(os.path.basename(path))
supersedes = records[number]["front"].get("superseded-by", "") supersedes = records[number]["front"].get("superseded-by", "")
@@ -241,13 +247,20 @@ def check_supersession_symmetry(failures, records):
failures.add("supersession", rel(record["path"]), f"superseder does not exist: {by}") failures.add("supersession", rel(record["path"]), f"superseder does not exist: {by}")
continue continue
other = records[match.group(1)] other = records[match.group(1)]
claims = os.path.basename(str(other["front"].get("supersedes", ""))) # `supersedes:` may name one record or several. One decision replacing two is a real
if claims != record["name"]: # situation -- two records that built and refined the same wrong mechanism are withdrawn
# by the one record that removes it -- and a check that allows only one would force
# either a chain of pro-forma records or an unmarked supersession.
claimed = other["front"].get("supersedes", "")
if isinstance(claimed, str):
claimed = [claimed] if claimed else []
claims = [os.path.basename(str(entry)) for entry in claimed]
if record["name"] not in claims:
failures.add( failures.add(
"supersession", "supersession",
rel(other["path"]), rel(other["path"]),
f"ADR {number} says this supersedes it; this record does not say so " f"ADR {number} says this supersedes it; this record does not say so "
f"(supersedes: {claims or 'absent'})", f"(supersedes: {', '.join(claims) or 'absent'})",
) )
@@ -299,8 +312,13 @@ def check_progressive_insights(failures, records):
unmarked change stands out as the anomaly it is. unmarked change stands out as the anomaly it is.
""" """
phrase = re.compile(r"progressive insight", re.I) phrase = re.compile(r"progressive insight", re.I)
marker = re.compile(r"\*\*Progressive insights?\s*[\u2014\u2013-]\s*(\d{4}-\d{2}-\d{2})\.?\*\*") # Both patterns stay on one line: a bold run does not span paragraphs, and `[^*]*` across
loose = re.compile(r"\*\*[^*]*[Pp]rogressive insights?[^*]*\*\*") # newlines will happily join an unrelated `**` far above to the marker below, reporting the
# whole span between them. It did exactly that the first time this ran.
# Trailing words after the date are allowed — "— 2026-09-26, correcting the one above." — so
# an insight can say what it relates to. Only the date's presence and position are fixed.
marker = re.compile(r"\*\*Progressive insights?[ \t]*[\u2014\u2013-][ \t]*(\d{4}-\d{2}-\d{2})[^*\n]*\*\*")
loose = re.compile(r"\*\*[^*\n]*[Pp]rogressive insights?[^*\n]*\*\*")
iso = re.compile(r"^\d{4}-\d{2}-\d{2}$") iso = re.compile(r"^\d{4}-\d{2}-\d{2}$")
for number, record in sorted(records.items()): for number, record in sorted(records.items()):
@@ -313,6 +331,10 @@ def check_progressive_insights(failures, records):
for m in loose.finditer(text): for m in loose.finditer(text):
if any(s <= m.start() and m.end() <= e for s, e, _ in good): if any(s <= m.start() and m.end() <= e for s, e, _ in good):
continue continue
# A bold run carrying a link is discussing an insight — usually another record's —
# rather than marking one. A marker never needs to cite anything.
if "](" in m.group(0):
continue
failures.add("insights", rel(record["path"]), failures.add("insights", rel(record["path"]),
"a progressive insight is not in the dated marked form " "a progressive insight is not in the dated marked form "
"'**Progressive insight \u2014 YYYY-MM-DD.**': %s" % m.group(0)) "'**Progressive insight \u2014 YYYY-MM-DD.**': %s" % m.group(0))
@@ -328,7 +350,15 @@ def check_progressive_insights(failures, records):
if any(s <= m.start() and m.end() <= e for s, e in covered): if any(s <= m.start() and m.end() <= e for s, e in covered):
continue continue
line = text.rfind("\n", 0, m.start()) + 1 line = text.rfind("\n", 0, m.start()) + 1
if text[line:m.start()].lstrip().startswith("#"): end = text.find("\n", m.end())
whole = text[line:end if end != -1 else len(text)]
if whole.lstrip().startswith("#"):
continue
# A line that also carries a link is discussing the rule, not marking a correction:
# a marker never needs to cite anything, and a record that reasons about the policy
# must be able to name it. Bare prose with no citation is the informal marking this
# is here to catch.
if "](" in whole:
continue continue
if loose.search(text, line, text.find("\n", m.end()) + 1 or len(text)): if loose.search(text, line, text.find("\n", m.end()) + 1 or len(text)):
continue continue
+17 -3
View File
@@ -31,8 +31,22 @@ another — and a mesh you cannot name precisely is a mesh two people describe d
- **store** — the one postgres server. It holds the controller's own context databases - **store** — the one postgres server. It holds the controller's own context databases
(`inventory`, `identity`, `licences` — a context owns its store, [ADR 0008](../02-DECISIONS/0008-a-context-owns-its-store.md)) (`inventory`, `identity`, `licences` — a context owns its store, [ADR 0008](../02-DECISIONS/0008-a-context-owns-its-store.md))
and every module's own database. One server, many databases — never one shared "mesh database". and every module's own database. One server, many databases — never one shared "mesh database".
- **broker** — the one lavinmq message bus. It carries the mesh bus on the `/` vhost and a vhost per - **bus** — the mesh's own nervous system: NATS, one per mesh, carrying every link the mesh has —
consumer that requires `amqp`. control, declarations, builds, events, tool calls
([ADR 0106](../02-DECISIONS/0106-the-bus-is-nats.md)). A module reaches it by requiring
`mesh-bus` ([ADR 0128](../02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md)); one that
does not require it has no account on it. Held by the `mesh-broker` seat, which is named after
the *role* rather than the server, so the server can change without the seat doing so.
- **the deprecated broker** — the lavinmq module. It was the mesh's bus and is not any more. It
keeps running as an **ordinary provider** of the `amqp` provision, for modules that need a
message broker of their own the way something needs a database
([ADR 0127](../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md))) — no seat, not foundation,
never raised at genesis, and a mesh that never installs it is complete.
Say *the deprecated broker*, not "the compatibility broker" (it serves the mesh's own modules,
not only the predecessor's) and not "the AMQP broker" (naming it after a protocol invites
describing the bus by contrast with it, which is backwards: the bus is the mesh's nervous
system and this is a module).
## What the mesh stores and serves ## What the mesh stores and serves
@@ -47,7 +61,7 @@ another — and a mesh you cannot name precisely is a mesh two people describe d
- **seat** — a named role at a scope (node / site / mesh), held by a module assignment, from a - **seat** — a named role at a scope (node / site / mesh), held by a module assignment, from a
**closed set** the mesh defines: a claim naming a seat outside the set is refused. A seat may **closed set** the mesh defines: a claim naming a seat outside the set is refused. A seat may
**deliver a provision**, and its holder is then the mesh's answer for it when several modules **deliver a provision**, and its holder is then the mesh's answer for it when several modules
provide it ([ADR 0110](../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md)). provide it ([ADR 0126](../02-DECISIONS/0126-a-module-declares-its-own-seats.md) (superseding [ADR 0110](../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md))).
The set, with who holds each seat, is the overview of what a mesh has The set, with who holds each seat, is the overview of what a mesh has
([26 — The seats](../03-DESIGN/01-to-be/26-the-seats.md)). A seat has a **capacity**: a ([26 — The seats](../03-DESIGN/01-to-be/26-the-seats.md)). A seat has a **capacity**: a
capacity-1 seat is exclusive (one holder); a higher-capacity seat is a **bench** (several holders capacity-1 seat is exclusive (one holder); a higher-capacity seat is a **bench** (several holders
+14 -1
View File
@@ -21,7 +21,12 @@ incident someone must **clear**.
## Steps ## Steps
1. Take the next free number. Create `04-ISSUES/NNN-short-name/00-report.md`: 1. Take the next free number — **across `main` and every open pull request**, not `main` alone.
Work sits on unmerged branches for days, so two people both reading `main` allocate the same
number; it happened twice in one hour between two machines, and the second collision reached
`main` with every check passing (issue 155). `cycle.py` now refuses two records sharing a number,
which catches a collision but does not prevent one. Create
`04-ISSUES/NNN-short-name/00-report.md`:
```yaml ```yaml
--- ---
@@ -42,6 +47,14 @@ incident someone must **clear**.
## Rules ## Rules
- Closed issues are never deleted — they are the mesh's symptom-to-component memory. - Closed issues are never deleted — they are the mesh's symptom-to-component memory.
- `fixed-by:` names something that will still exist: a commit or a pull request, never a branch. A
branch is deleted when it merges, so a branch name there is a pointer that resolves to nothing by
the time anybody follows it.
- A fix that turns out to have broken something else is written back into the record that asked for
it, pointing at the new issue. Somebody arriving at a record to learn why the code is the way it
is must not have to already know there was a sequel.
- Renumbering a collision happens once, in the branch that lands last. Renumbering a branch whose
author is still pushing only moves the race.
- An issue whose answer is a general lesson should also be written to the knowledge base, so - An issue whose answer is a general lesson should also be written to the knowledge base, so
the next person searching a symptom finds it. Both, not either. the next person searching a symptom finds it. Both, not either.
- `status: wontfix` is legitimate and requires a sentence saying why. - `status: wontfix` is legitimate and requires a sentence saying why.
+8
View File
@@ -13,6 +13,14 @@ decisions taken over three days; the reasoning is kept, the fragmentation is not
The environment a change is run against before it reaches real machines. The environment a change is run against before it reaches real machines.
> **Still the lab, no longer the test bed — 2026-09-30, by [ADR 0149](0149-the-live-mesh-is-the-test-bed.md).**
> Everything here stands. What changed is what the lab is *for*: a change is verified against the mesh
> that is running, because the faults that cost the most are faults of a mesh that already exists —
> bound consumers, containers made against an older roster, an adopted machine — and a bed is by
> construction a mesh that does not. Raising a mesh from bare is now the lab's whole job, which is the
> one thing the live mesh cannot be asked to do. 0149 also supersedes
> [ADR 0068](0068-the-lab-takes-requests.md), which extended this one and was never built.
## A node in the lab is a virtual machine ## A node in the lab is a virtual machine
It boots a stock Linux image, runs the real install, and becomes a node. **It is not a model of It boots a stock Linux image, runs the real install, and becomes a node. **It is not a model of
+17 -1
View File
@@ -1,6 +1,6 @@
--- ---
topic: building it topic: building it
status: proposed status: accepted
date: 2026-09-01 date: 2026-09-01
deciders: jochen deciders: jochen
reconstructed: false reconstructed: false
@@ -101,3 +101,19 @@ the digest down after building.
**Whether kind 4 deserves a module at all.** Thirty-five descriptions that say *install this and **Whether kind 4 deserves a module at all.** Thirty-five descriptions that say *install this and
write these files* may be better as one module with settings than as thirty-five modules. Left write these files* may be better as one module with settings than as thirty-five modules. Left
open deliberately; it is a question about the shape of the catalogue, not about whether to have one. open deliberately; it is a question about the shape of the catalogue, not about whether to have one.
## Accepted, 2026-09-29, against what was built
*Marked in a grooming pass: the mesh was built to this record and the record still said `proposed`.*
The proposal is the arrangement that exists. `mesh-catalog` holds descriptions of software we did
not write and the programs that provision it, and holds neither the mesh's own components nor an
application's own module. The mesh's list of modules is a table in the control plane, filled by
`module add`, and every module records the source it came from with the commit it was read at.
**One half is not built: `module check` as a command on the control plane's binary.** A manifest is
still validated by a test that reaches into the control plane's internals — which works for this
catalogue and gives nothing at all to somebody describing their own application in their own
repository, which this record says is the case that matters most. That is
[issue 148](../04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md).
@@ -76,6 +76,22 @@ three relationships, one broker, one runtime, all declared on the manifest.
- The runtime must dispatch a module's event handlers as well as its tools; that generalisation is - The runtime must dispatch a module's event handlers as well as its tools; that generalisation is
small (both arrive by importing the module's entrypoint) but it is real work. small (both arrive by importing the module's entrypoint) but it is real work.
## Progressive insight
> **Progressive insight — 2026-09-26.** *"No provisioner and no per-consumer setup" was a fact
> about the transport, and the transport changed.* This record's table says an event's machinery is
> "nothing but the broker's topic routing", and the text that an event needs "no per-consumer setup
> — only a subscription". That was true of a topic exchange, where a binding cost nothing and the
> broker fanned out. On NATS
> ([ADR 0106](0106-the-bus-is-nats.md)) a subscription is a **durable consumer**: a real object
> with a name, an ack policy, a delivery limit and its own ack subject, created when a module is
> assigned and removed when it is not. Per-consumer setup exists, and the controller does it.
>
> The decision is untouched — events are declared on both sides, 1:many, credential-free, and
> still provisioning's lighter sibling; the lightness is now relative rather than absolute.
> [ADR 0126](0126-a-module-declares-its-own-seats.md) adds the relationship this record's two
> columns had no room for: work addressed to a role, where exactly one holder must act.
## References ## References
- [ADR 0002](0002-nodes-communicate-over-a-broker.md) — the broker events ride. - [ADR 0002](0002-nodes-communicate-over-a-broker.md) — the broker events ride.
@@ -26,6 +26,14 @@ runtime is per-module, not per-node, and treating the audit-logger as special le
modules' code with nothing to run it: the conversion produced tools and events that, as it stands, modules' code with nothing to run it: the conversion produced tools and events that, as it stands,
never execute. never execute.
> **The hosting form is settled elsewhere — 2026-09-30.** Where this record says "a container", read
> [ADR 0150](0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md): a module's own
> code runs as supervised processes under this record's one account. Nothing else here changes — the
> per-module runtime, the per-tool key and the single scoped account are the argument this record made
> and they are why 0150 goes the way it does. The note is here because two design documents chose the
> other form without knowing this record existed
> ([issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md)).
## Decision ## Decision
### A module with tools or events runs a process of its own ### A module with tools or events runs a process of its own
@@ -80,6 +80,15 @@ reaching the routed name, which the clause above has just made resolvable inside
three are one decision: **compose the name, propagate it, certify it** — each is meaningless without three are one decision: **compose the name, propagate it, certify it** — each is meaningless without
the one before it. the one before it.
> **The mechanism changed — 2026-09-30, by [ADR 0148](0148-the-meshs-names-are-resolved-not-copied-into-containers.md).**
> A routed name still reaches every asker in the mesh, which is what this record decided and it stands.
> It no longer reaches them by being written into each declared container: copying the roster in made the
> roster part of every container's identity, so one name moving replaced every container in the mesh
> ([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). A
> container resolves through its machine's resolver instead. The consequence below — that an internal
> issuer's challenge needs the routed name resolvable inside the mesh — holds unchanged, by the means the
> machine itself already uses.
## Consequences ## Consequences
- **Lab-versus-production is one node-level `public-domain` setting**, not an override on every - **Lab-versus-production is one node-level `public-domain` setting**, not an override on every
+8 -1
View File
@@ -1,14 +1,21 @@
--- ---
topic: building it topic: building it
status: proposed status: superseded
date: 2026-09-12 date: 2026-09-12
deciders: jochen deciders: jochen
reconstructed: false reconstructed: false
extends: 0016-the-lab.md extends: 0016-the-lab.md
superseded-by: 02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md
--- ---
# 68. The lab takes requests, one at a time, and runs each from its own copy # 68. The lab takes requests, one at a time, and runs each from its own copy
> **Superseded — 2026-09-30, by [ADR 0149](0149-the-live-mesh-is-the-test-bed.md).** Never built. The
> live mesh became the test bed, because the faults that cost the most are faults of a mesh that
> already exists — bound consumers, containers made against an older roster, an adopted machine — and a
> bed is by construction a mesh that does not. The rule worth keeping from below is that a run reads a
> copy that is not anybody's working tree.
## Context ## Context
**The lab is exclusive hardware, and today a person holds it.** Raising a scenario takes over **The lab is exclusive hardware, and today a person holds it.** Raising a scenario takes over
@@ -35,6 +35,16 @@ The development cycle is enforced mechanically, to the extent frontmatter can ca
capability existed" is an answer). capability existed" is an answer).
- **No silent graduation** — a `graduated` research overview says what it `became:`, and the - **No silent graduation** — a `graduated` research overview says what it `became:`, and the
targets exist. targets exist.
- **No two records answering to one number** — added 2026-09-30; see the insight below.
> **Progressive insight — 2026-09-30.** The list above named four things `cycle.py` enforces, and
> now names five. Nothing enforced that two issue records hold different numbers: two machines
> filing issues within one hour both read `main`, both took "the next free number", and collided
> twice — the second collision reaching `main` with `records.py`, `cycle.py` and `index.py` all
> reporting success (04-ISSUES/155). A number is how every other record cites one, so two records
> answering to it is a citation that resolves to whichever folder the reader opened. `cycle.py`
> refuses it now. The decision here stands exactly as written: this is one more thing frontmatter
> and file names can carry, found by its absence rather than by reasoning.
[`00-META/checks/cycle.py`](../00-META/checks/cycle.py) refuses violations, beside `records.py` [`00-META/checks/cycle.py`](../00-META/checks/cycle.py) refuses violations, beside `records.py`
and `index.py`; all three run before any merge here. What frontmatter cannot see — that code work and `index.py`; all three run before any merge here. What frontmatter cannot see — that code work
+24
View File
@@ -75,6 +75,30 @@ after its deliveries are exhausted; a module's account cannot publish outside it
subscribe outside its `consumes`. Then the cutover bed: a mesh on AMQP with the predecessor's subscribe outside its `consumes`. Then the cutover bed: a mesh on AMQP with the predecessor's
compatibility broker beside it moves its bus in one rollout with every node reporting afterwards. compatibility broker beside it moves its bus in one rollout with every node reporting afterwards.
## Progressive insight
> **Progressive insight — 2026-09-26.** *The compatibility broker was not single-purpose when this
> was written.* This record says the adopted AMQP broker is "kept as a module with one purpose —
> the predecessor's clients". Two modules of the new mesh also depended on it, through a `requires:
> ["amqp"]` grant its provisioner answered with a private vhost — `amqp-ping` and
> `amqp-email-forwarder`. On the retirement condition below, both would have been left requiring
> something no provider answers.
> [ADR 0125](0125-the-bus-is-the-only-broker.md) resolves it by moving them onto the bus and
> retiring the interface, which makes this record's sentence true rather than merely intended. The
> decision — the bus is NATS, the AMQP broker becomes the predecessor's compatibility broker and
> retires with the last of them — is unchanged.
> **Progressive insight — 2026-09-26, correcting the one above.** *The broker is not a
> compatibility module at all, and the sentence does not become true.* The insight above said
> [ADR 0125](0125-the-bus-is-the-only-broker.md) would make "one purpose — the predecessor's
> clients" true by moving the mesh's own modules off it.
> [ADR 0127](0127-amqp-is-a-provision-not-the-bus.md) supersedes that: nothing moves off, because
> a module may legitimately need an AMQP broker as a backing service the way it needs a database.
> The broker becomes **an ordinary provider module** — no seat, not foundation, not raised at
> genesis, and with no retirement condition, because the day its last client disappears is not a
> day anything is waiting for. What this record decided — the mesh's bus is NATS — is untouched
> by both; what was wrong was the sentence describing what happens to the old server, twice.
## References ## References
- [research 014](../01-RESEARCH/014-the-bus-on-nats/00-overview.md) - [research 014](../01-RESEARCH/014-the-bus-on-nats/00-overview.md)
@@ -9,6 +9,17 @@ extends: 0009-modules-and-the-graph.md
# 110. A seat is held by one assignment, from a closed set, and it may deliver a provision # 110. A seat is held by one assignment, from a closed set, and it may deliver a provision
> **Narrowed, not replaced — 2026-09-27, on merging two lines of work.** This was marked superseded by
> [ADR 0126](0126-a-module-declares-its-own-seats.md), and that overstated it: 0126 says in as many
> words that *"everything 0110 decided about what a seat is stands untouched"*. What moved is where the
> set lives and who may add to it —
> [0126](0126-a-module-declares-its-own-seats.md) lets a module declare one and makes the set derived,
> [0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md) names the mesh's own
> for their scope, and [0122](0122-a-seat-is-data-a-rename-is-a-database-update.md) moves them out of
> code into a table. **What a seat *is* — one holder at its scope, a definition saying what a module
> can hold against an assignment saying what it does, a role made singular rather than a module — is
> this record and still current**, which is why those three rest on it.
## Context ## Context
[ADR 0009](0009-modules-and-the-graph.md) introduced claims: a module declares something [ADR 0009](0009-modules-and-the-graph.md) introduced claims: a module declares something
@@ -1,6 +1,6 @@
--- ---
topic: what runs on it topic: what runs on it
status: proposed status: accepted
date: 2026-09-25 date: 2026-09-25
deciders: jochen deciders: jochen
reconstructed: false reconstructed: false
@@ -1,6 +1,6 @@
--- ---
topic: what runs on it topic: what runs on it
status: proposed status: accepted
date: 2026-09-25 date: 2026-09-25
deciders: jochen deciders: jochen
reconstructed: false reconstructed: false
@@ -275,3 +275,11 @@ On acceptance, each of these is superseded or amended by this record, not edited
everything a module needs is a requirement everything a module needs is a requirement
- [Issue 095](../04-ISSUES/095-a-module-assigned-after-genesis-has-no-broker-account/00-report.md), - [Issue 095](../04-ISSUES/095-a-module-assigned-after-genesis-has-no-broker-account/00-report.md),
[issue 103](../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md): what fails today [issue 103](../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md): what fails today
## Accepted, 2026-09-29, against what was built
*Marked in a grooming pass.* The vault is a module providing `secret` at mesh scope, and six
modules in the catalogue require it — so a shared secret is a requirement answered by the vault,
which is what this record asks for. Private keys are still made where they are used and never
travel, which is the other half and was never in question.
@@ -1,6 +1,6 @@
--- ---
topic: what runs on it topic: what runs on it
status: proposed status: accepted
date: 2026-09-26 date: 2026-09-26
deciders: jochen deciders: jochen
reconstructed: false reconstructed: false
@@ -189,6 +189,23 @@ On acceptance, each of these is amended by this record, not edited:
credential is no longer all-or-nothing with a window. It overlaps, with each step confirmed. credential is no longer all-or-nothing with a window. It overlaps, with each step confirmed.
A single-party credential is staged, not replaced. A single-party credential is staged, not replaced.
## Accepted, 2026-09-30, and not scheduled
Accepted as written. The separation it draws — a consumer's *resource* and a *credential that reaches
it* are different things with different lifecycles — is the part that had to be settled, because the
alternative is what the record was written against: retiring a credential taking the data it reached
with it. That is a data-loss shape, and a record that names it should not sit unresolved while the
code that could hit it is being written.
**It is not built, and accepting it does not schedule it.** The SDK's provisioner adapter is still
`create` / `remove` / `holds` rather than the four operations above, and no provider implements the
two-credential rotation. Accepted-and-not-built is an ordinary state here — 0141 and 0142 are both in
it — and it is the honest one: leaving this `proposed` made it invisible to anyone reading what the
mesh has decided, while changing nothing about what runs.
The work it implies belongs with the provisioner contract, beside
[issue 124](../04-ISSUES/124-a-consumer-cannot-be-told-what-its-provider-derived/00-report.md).
## Consequences ## Consequences
- **Every credential provider's adapter changes**, in two steps. The first separates *retire a - **Every credential provider's adapter changes**, in two steps. The first separates *retire a
@@ -1,6 +1,6 @@
--- ---
topic: what runs on it topic: what runs on it
status: proposed status: accepted
date: 2026-09-26 date: 2026-09-26
deciders: jochen deciders: jochen
extends: 0112-a-module-definition-names-no-node-mesh-or-path.md extends: 0112-a-module-definition-names-no-node-mesh-or-path.md
@@ -47,3 +47,10 @@ other boundary already is: the module name.
node runs one of each (ADR 0115)" — instead of failing on whichever name collides first. node runs one of each (ADR 0115)" — instead of failing on whichever name collides first.
- Multi-tenant asks are answered in the catalogue (a second module definition), not in the - Multi-tenant asks are answered in the catalogue (a second module definition), not in the
control plane. control plane.
## Accepted, 2026-09-29, against what was built
*Marked in a grooming pass.* The rule is enforced where it cannot be forgotten: `assignment`'s
primary key is `(node, module)`, so a second assignment of one module to one machine is not a thing
the mesh can hold. The record read `proposed` while the schema had already settled it.
@@ -12,7 +12,7 @@ extends: 0102-the-mesh-writes-into-a-shared-file-never-over-it.md
## Context ## Context
When a resource stops being declared — its module unassigned, the node sent a When a resource stops being declared — its module unassigned, the node sent a
deliberately-empty declaration ([issue 127](../04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)), deliberately-empty declaration ([issue 149](../04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)),
or a new catalogue version renaming its id — the host undoes it. The host's own code states or a new catalogue version renaming its id — the host undoes it. The host's own code states
the rule it means to follow: **it removes what it made and leaves what it merely configured.** the rule it means to follow: **it removes what it made and leaves what it merely configured.**
For almost every resource it does exactly that: For almost every resource it does exactly that:
@@ -40,7 +40,7 @@ undeclare can do. Found reviewing the uplink modules
- the uplink modules would have stopped the network manager, taking the machine off the only - the uplink modules would have stopped the network manager, taking the machine off the only
link the mesh reaches it by. link the mesh reaches it by.
[ADR 0117](0117-a-machines-uplink-is-a-seat.md) answered that for its own modules with a [ADR 0125](0117-a-machines-uplink-is-a-seat.md) answered that for its own modules with a
service declared with no `state`. Every other module that declares a unit it did not make is service declared with no `state`. Every other module that declares a unit it did not make is
exposed in the same way, and relying on each author to remember an opt-out is how the next one exposed in the same way, and relying on each author to remember an opt-out is how the next one
is missed. is missed.
@@ -91,7 +91,7 @@ records the state it first found the unit in, and undeclaring returns the unit t
- The service's settings the mesh wrote are given back by their own resources (a kept original - The service's settings the mesh wrote are given back by their own resources (a kept original
restored, a region or keys removed). A running service keeps running on what it read until it restored, a region or keys removed). A running service keeps running on what it read until it
next reads its configuration; the mesh does not restart it to make it notice. next reads its configuration; the mesh does not restart it to make it notice.
- A service declared with no `state` (ADR 0117) remains the way to say the mesh must not - A service declared with no `state` (ADR 0125) remains the way to say the mesh must not
**start** a unit either; undeclared, it is forgotten. **start** a unit either; undeclared, it is forgotten.
## Consequences ## Consequences
@@ -113,7 +113,7 @@ records the state it first found the unit in, and undeclaring returns the unit t
## References ## References
- [issue 130](../04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md): the finding - [issue 130](../04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md): the finding
- [ADR 0117](0117-a-machines-uplink-is-a-seat.md): the uplink modules, and a service with no state - [ADR 0125](0117-a-machines-uplink-is-a-seat.md): the uplink modules, and a service with no state
- [ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md), [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md): - [ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md), [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md):
what is given back, and how what is given back, and how
- mesh-host `internal/apply/apply.go` (`remove`, the service case) - mesh-host `internal/apply/apply.go` (`remove`, the service case)
@@ -53,7 +53,7 @@ is removed from where its unit reads it.**
reporting it. reporting it.
- **The mesh never brings it back.** Undeclaring the private network does not restore the found - **The mesh never brings it back.** Undeclaring the private network does not restore the found
tunnel: the mesh stopped it, and nothing is started on the way out tunnel: the mesh stopped it, and nothing is started on the way out
([ADR 0118](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md)). A machine whose ([ADR 0126](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md)). A machine whose
private network is unassigned has no tunnel until it is assigned again — which is what private network is unassigned has no tunnel until it is assigned again — which is what
unassigning it means. unassigning it means.
@@ -83,5 +83,5 @@ is removed from where its unit reads it.**
- [ADR 0105](0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md): the take, and why it keeps - [ADR 0105](0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md): the take, and why it keeps
the found configuration during it the found configuration during it
- [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md): kept originals - [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md): kept originals
- [ADR 0118](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md): nothing is started on the way out - [ADR 0126](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md): nothing is started on the way out
- mesh-host `internal/apply/takeover.go` - mesh-host `internal/apply/takeover.go`
@@ -0,0 +1,136 @@
---
topic: what runs on it
status: accepted
date: 2026-09-27
deciders: jochen
reconstructed: false
extends: 0112-a-module-definition-names-no-node-mesh-or-path.md
---
# 120. A roster fact carries its format as a template: the mesh owns the data, the module owns the format
## Context
A **fact** is a thing only the mesh knows — which machines exist, what they are called, where they
are — written into a file where a module asks for it. The mesh computes it from the graph; a module
loads it, restarts on it, does what its software does with it. Facts replaced three modules that
existed only because computed output needed somewhere to live and ran no software of their own
([ADR 0040](0040-what-a-module-is.md)).
But the *format* lived in the control plane. A fact was a name from a closed list, and each name had
a formatter written in Go beside the others: `node-names` wrote the roster as an `/etc/hosts` file,
`node-zones` wrote it as a dnsmasq resolver's `local=`/`address=` lines. Adding a consumer meant
adding a formatter — in the consumer's own configuration language — to the mesh.
The ssh work made the cost plain. An operator's `~/.ssh` wants three roster projections — a
`known_hosts`, an ssh `config` of `Host` blocks, an `authorized_keys` — each in ssh's syntax. Under
the closed list that is three more formatters in the control plane, teaching it ssh's configuration
language. And it does not stop at ssh: every daemon that reads the roster in its own file format
would put its grammar here. The control plane was accreting the configuration languages of software
it does not run — the exact thing [ADR 0040](0040-what-a-module-is.md) says is a
module's and not the mesh's.
The shape underneath is one shape. WireGuard's `[Peer]` blocks, `/etc/hosts`, dnsmasq's zones, an
ssh `known_hosts` — all of them are *the roster, projected into a file*. Only the projection differs,
and the projection belongs to whoever runs the software that reads it.
## Considered Options
**1. Keep the closed list; add a formatter per consumer.** Rejected. The control plane learns the
configuration language of every daemon any module might run, without bound, and each format lives in
the mesh rather than in the module that owns the file. A module cannot change how its own file is
written without a control-plane change.
**2. A general placeholder vocabulary over `content`, like `${machine:address}` but for the
roster.** Rejected. The mesh's other substitutions each resolve to *one* scalar — this machine's
address, one provider's port. The roster is inherently a *repetition*: one block per machine. A flat
`${…}` vocabulary cannot iterate, and a mechanism that could would be a template in all but name.
**3. The module gives a path and a template over the roster; the mesh renders it.** Chosen. The mesh
owns the data — who exists, their names and addresses — and hands it to a Go `text/template` the
module wrote. The mesh renders and reads neither the template's intent nor the file's meaning.
## Decision
**A fact is a path and a template.** In a module's manifest, `facts` maps a name the module chooses
to a `{ path, template }`. The template is a Go `text/template` over a fixed **roster view**:
- `.Node` — this machine's bare name.
- `.Suffix` — what a mesh name ends in (`internal`, or the operator's choice), as composed.
- `.Names` — every name the mesh serves: the machines *and* the names it was told to route.
- `.Machines` — only the machines that are nodes of this mesh.
Each of `.Names` and `.Machines` is a list of `{ Name, FQDN, Address }`. A machine the mesh has a
record for but cannot yet place has no address and is left out of both — a name that resolves to
nothing is a connection that hangs, so it is omitted rather than written (the same rule as before).
**The mesh owns the data; the module owns the format.** The control plane holds **no** formatter.
The two built-in projections render through the same path any module uses:
- **`/etc/hosts`** is a template on the mesh's own network module. The mesh writes `/etc/hosts`
because being on the private network is what gives a machine a name — but the *layout* is a
template like any other, shipped with the control plane because that module ships with it, not
because the control plane knows the hosts-file format.
- **dnsmasq's zones** move into dnsmasq. The `local=`/`address=` grammar is dnsmasq's configuration
language, and it now lives in dnsmasq's manifest, where the module that runs dnsmasq owns it.
**A fact also says whether its file is the mesh's whole or a region of the machine's.** A hosts file
is the machine's — its `localhost`, the operator's lines, another tool's marked blocks — so
`node-names` is `shared`: the mesh owns only its region and keeps the rest byte for byte, the host
laying it down `into: block`
([ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md), hq issue 128). A resolver's
zones file is the mesh's whole, and is not shared. The template renders the content either way;
`shared` decides how the host writes it. This composes with hq 128 rather than replacing it: the
region *mechanism* is the host's, the region's *format* is the module's template.
**The names-vs-machines distinction is the template's choice** ([04-ISSUES/111](../04-ISSUES/111-the-resolver-is-told-names-the-mesh-serves-not-only-machines/00-report.md)):
a container's hosts ranges `.Names`, so a routed name resolves to the machine serving it; a resolver
told the mesh's suffix is its own ranges `.Machines`, or a routed name written there with the suffix
appended is a name nobody will ever ask for.
**A template that will not render is refused at composition, not on a machine.** A template that does
not parse, or reads a field the roster does not have, fails where the manifest is — the closed-list
safety, moved from the fact's *name* to the roster's *shape*. A daemon that starts, reads a file the
mesh could not render, and answers nothing is a much worse way to find out.
**WireGuard stays a computed generator, and that is the line.** Its `mesh0.conf` is not a pure roster
projection — it carries topology the control plane decides: which peers are reachable, endpoints, hub
forwarding, keepalive for a NAT'd node. And it is *foundational*: the overlay must be up before any
module can be delivered, so the thing that writes it cannot itself be a delivered module. The line
this draws: **the substrate that delivery rides on is the control plane's; everything layered on a
working overlay is a roster template.** DNS, hosts, and ssh are layered; the overlay is the floor.
## Consequences
- **ssh is two templates and no control-plane change.** Once the roster view carries a machine's ssh
host key and its operator account ([to-be 29](../03-DESIGN/01-to-be/29-a-node-has-operator-accounts.md)),
`known_hosts`, the ssh `config`, and `authorized_keys` are templates on the ssh modules — the mesh
gains no knowledge of ssh's syntax. This ADR is what makes that work land without touching the
controller.
- **A new roster projection never touches the control plane.** Any module that reads the roster in
its own format ships its own template.
- **A module can change how its own file is written** without a control-plane change — it is editing
its own manifest.
- **The schema changed and is not backward compatible.** A fact was a string (a path); it is now
`{ path, template }`. The old string form has no template and cannot be auto-upgraded, because the
format it implied was the formatter this ADR deletes. The controller and every catalogue module
using facts — only dnsmasq — land together. A controller and a catalogue that disagree cannot
compose the module: the running daemon on a machine is unaffected, but the mesh will not send it a
new declaration until both sides agree.
- **The output did not change.** The `/etc/hosts` and dnsmasq zones a machine receives are
byte-for-byte what the deleted formatters wrote, pinned by tests that render the built-in template
and compose the real dnsmasq manifest.
## References
- [ADR 0040](0040-what-a-module-is.md): a module is software the mesh runs — a
format the mesh knows for software it does not run was the accretion this stops
- [ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md): a module definition names no
path; this is its sibling for content — a module definition names no format the mesh must know
- [to-be 29](../03-DESIGN/01-to-be/29-a-node-has-operator-accounts.md): the ssh consumer this
unblocks, and the roster fields it will add
- [04-ISSUES/111](../04-ISSUES/111-the-resolver-is-told-names-the-mesh-serves-not-only-machines/00-report.md): every served name is not a
machine — now the template's choice of `.Names` or `.Machines`
- mesh-controller `internal/catalogue/roster.go` (the mechanism), `internal/overlay/generator.go`
(the built-in `/etc/hosts` template), `internal/catalogue/manifest.go` (`RosterFile`)
- mesh-catalog `modules/dnsmasq/module.json` (the zones template, dnsmasq's own)
@@ -0,0 +1,132 @@
---
topic: what runs on it
status: accepted
date: 2026-09-27
deciders: jochen
reconstructed: false
extends: 0110-a-seat-is-a-module-assignment-from-a-closed-set.md
---
# 121. A system seat is named for its scope, and a module may define its own
## Context
[ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) made seats a closed set the
control plane defines: a well-formed name no longer becomes a seat by being claimed, so a person can
read what a mesh can have and who fills each role. It left two things unsettled that the growing set
now exposes:
- **The names carry no rule.** `mesh-controller`, `mesh-store`, `mesh-broker` are named for the mesh;
beside them sit `the-artifact-store`, `the-build-machine`, `the-dns-port`, `the-showcase`,
`the-uplink` — a second naming style with no principle behind it. A reader cannot tell a seat's
scope from its name, and the mesh's own roles do not look like the mesh's.
- **The set is the *only* place a seat may be defined.** A module claiming any name not in the
control plane's set is refused. That is right for *system* roles — one broker, one packet filter
per node — but it means a module can never define a role of its own: a demo module's
`the-showcase`, a future application's coordination role, must be smuggled into the control plane's
set or not exist. The control plane ends up holding roles that are not the mesh's to define.
Reviewing the set against these also found seats whose *scope* or *membership* is wrong, not just
their name — the review is the occasion to fix those too.
## Decision
**A system seat — one the control plane defines — is named for its scope:**
- **`mesh-*`** for a mesh-scoped seat: one holder in the whole mesh, a role the mesh has once
(`mesh-controller`, `mesh-store`, `mesh-broker`, `mesh-git`, …). A `mesh-*` seat is always held by
a module **on a named node** — `mesh-git` is gitea *on novox*, not "gitea"; another node running
gitea does not hold `mesh-git` unless it is the holder. The seat is the mesh's single answer for
the role, and which node answers is part of what the seat records.
- **`node-*`** for a node-scoped seat: one holder per node, a role each machine has at most once
(`node-packet-filter`, `node-intrusion-prevention`, `node-uplink`, …).
The three already-`mesh-*` seats keep their names; the rest are renamed by this rule. The scope a
name declares must match the seat's actual scope — a `mesh-*` seat at node scope, or the reverse, is
a contradiction the reader is entitled to trust is impossible.
**The control plane defines only system seats. A module may define its own.** A seat named `mesh-*`
or `node-*` is the control plane's, and claiming one the control plane does not define is refused as
before. Any *other* name is a **module-defined seat**: valid when the module declaring the claim also
declares the seat (its name, scope, and — if any — the protocol its holder speaks). The control plane
enforces one-holder-per-scope for it exactly as for its own, but does not otherwise know what it
means. So an application can coordinate its own instances through a seat of its own, and the mesh's
closed set stays what its name says it is: the *system's* roles, not everyone's.
**Specific seats this settles:**
- **`the-build-machine` → `mesh-build-machine`, and its scope becomes mesh.** There is one build
machine in the mesh (the builder on novox), not one per node. Node scope said the opposite. It
delivers no provision; it is the mesh's single build machine.
- **`the-private-network` → `mesh-private-network`, held by the network *server* on one node.** Today
it is node-scoped and held on every node, with a stated (untested) story that a different VPN could
hold it per machine — which would force every provider module to independently implement receiving
and applying the controller-composed configuration. The mesh does not work that way and should not
pretend to: **one mesh decides one private network.** The seat is mesh-scoped, held by the server
module (WireGuard on the hub, novox). A machine that joins is given a **client module** that
receives the composed configuration and applies it; when a node joins, the mesh emits each node's
configuration so all of them know each other at once. This drops per-node VPN choice deliberately —
the private network is nox-mesh's own, and it defines the nodes' configuration rather than being
assembled from each node's opinion. (Implementation: the overlay generator's per-node computation
is unchanged; what changes is the seat's scope and the server/client split of the module.)
- **`the-showcase` → removed from the set; it becomes a module-defined seat.** It is a demo module's
own coordination role, claimed by nothing else and held nowhere. It is the first module-defined
seat, and the reason the rule above is needed rather than hypothetical.
- **`the-dns-port` → `node-dns-resolver`** (the daemon that binds `:53`), kept distinct from
**`the-resolver-configuration` → `node-resolver-config`** (what writes `resolv.conf`). Two roles,
two seats; the rename must not blur them.
- **`the-packet-filter` → `node-packet-filter`** and **`the-intrusion-prevention` →
`node-intrusion-prevention`** — names kept as-is but for the prefix. "Packet filter" stays distinct
from "firewall", which would swallow intrusion-prevention too.
- **`the-uplink` → `node-uplink`** ([ADR 0125](0117-a-machines-uplink-is-a-seat.md)). Unheld, so it
renames with no migration.
- **The registry seats — `the-artifact-store`, `npm-package-registry` (→ `mesh-artifact-store`,
`mesh-npm-package-registry`) — and `git` (→ `mesh-git`) — are decided but deferred.** They each
*deliver* a provision, so renaming them is a delivering-seat migration: a holder that stops
resolving mid-flight takes a provision away from every consumer. That risk is not worth carrying in
the same pass as the node-* renames, so they keep their names until done deliberately.
**`distribution` stays the mesh's registry; only `verdaccio` is retired.** An earlier draft of this
record had the registry consolidating onto gitea and `distribution` retired — that was reversed:
`distribution` is the standalone OCI registry serving every `artifact-store://…@sha256` image (the
control plane's own included), and the mesh keeps it. `verdaccio` was a *second* npm registry;
gitea already provides `npm-package-registry`, so verdaccio is redundant and is removed. It is only
in the catalogue (never registered in the running mesh), so removing it is deleting the module — no
migration, nothing to strand.
## Consequences
- **A reader learns a seat's scope from its name.** `mesh-*` is mesh-wide and one; `node-*` is
per-machine. The mesh's own roles finally look like the mesh's.
- **Applications get their own seats** without the control plane learning their meaning. The closed
set shrinks to what it should be — the system's roles — and stops being where unrelated roles hide.
- **The renames are a coordinated migration, not a rename.** A held seat's name lives in three places
that must move together: the control plane's set (`seats.go`), every claiming manifest, and what
each node reports it holds (re-derived by re-registering the manifest and re-pushing). A seat
renamed in one place and not the others stops resolving to its holder — and for a *delivering* seat
(`mesh-store`→postgres, `mesh-broker`→amqp, the registry seats) that is a mesh-wide provision
outage, the same failure mode as a schema change hitting an old manifest. So: the non-delivering
`node-*` seats and `mesh-build-machine` migrate as one tested controller+catalogue change;
`node-uplink` is free (unheld); the delivering registry seats are deferred to their own pass.
- **The node-* migration was done as one controlled step, and it froze briefly.** Deploying the new
controller made it reject the still-old-named claims in the stored manifests, so composition stopped
for the affected nodes until each manifest was re-registered under its new name; running services
were untouched, and the window was seconds. This is the coordinated-migration cost named above,
paid once — and the reason the *delivering* registry seats, whose freeze would be a provision
outage rather than a compose pause, are not folded into the same pass.
- **`distribution` is not retired.** It stays as the registry; only `verdaccio` (a redundant second
npm registry) is removed. The mesh keeps one OCI registry (`distribution`) and gitea for npm/git —
the "one registry, on gitea" idea was considered and dropped.
- **The private network stops pretending to be swappable per node.** The gain is a coherent
server/client model matching how the controller already composes configuration; the cost is that
choosing a different VPN is now a mesh-wide change, not a per-node one — accepted.
## References
- [ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) — the closed set this refines
- [ADR 0125](0117-a-machines-uplink-is-a-seat.md) — `the-uplink`, renamed here to `node-uplink`
- [ADR 0079](0079-the-foundation-seats-are-named-after-their-servers.md) — the original `mesh-*` seats
whose naming this generalises
- [to-be 26](../03-DESIGN/01-to-be/26-the-seats.md) — the seat table, updated by this
- mesh-controller `internal/catalogue/seats.go` (the set and claim validation),
`internal/overlay/generator.go` (the private network as server + client)
@@ -0,0 +1,107 @@
---
topic: what runs on it
status: accepted
date: 2026-09-27
deciders: jochen
reconstructed: false
supersedes-in-part:
- 0110-a-seat-is-a-module-assignment-from-a-closed-set.md
- 0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md
---
# 122. A seat is data the controller owns, and a rename is a database update
## Context
[ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) made the seats a closed set the
control plane defines, and [ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md)
named them by scope. Both were right about *what* a seat is. Both left it defined the wrong *way*:
**the set is a hardcoded Go slice compiled into the controller, and everything references a seat by
its name as a string literal.** Renaming `the-packet-filter` to `node-packet-filter` this session
took, in one pass:
- an edit to the Go slice in `internal/catalogue/seats.go`, recompiled into a new controller image;
- an edit to a `const gitSeat = "git"` in *production* control-plane code (`source.go`), because a
seat's name was hardcoded where a repository's home is resolved;
- edits to every claiming manifest in the catalogue, each re-registered;
- a controller **rebuild and redeploy**, which — because the running controller then refused the
still-old-named claims in stored manifests — **froze composition** for the affected nodes until
each manifest was re-registered under its new name;
- the same coupling in the **build machine**, which embeds the same seat set and refused to build
anything claiming a name it did not yet know;
- a **deadlock** when the build machine's own seat was renamed, since the old builder could not
build the new builder whose manifest claimed a name it rejected.
None of that is what a rename should cost. A rename is the operator changing a label. It should be a
single write, and nothing should have to be rebuilt, refused, or unfrozen. The set being *closed*
(0110) and *named by scope* (0121) are good rules; **the set being code is the mistake.** When
adhering to the design means twenty steps and a `const` in the resolver, the design is what to fix.
## Decision
**The seat set is data the control plane owns, not code it is compiled from.** The seats live in a
table in the controller's store — one row per seat: a **stable id**, a `name`, a `scope`, what it
`delivers` (a provision, or nothing), and the record that decided it. The rows are seeded by a
migration (the closed set 0110 defines still ships with the mesh), and thereafter they are ordinary
data the control plane reads and writes.
**A seat is referenced by its stable id, never by its name.** A claim, a held-seat record, and any
control-plane code that must name a seat (the git-seat resolver, the artifact-store guard) hold the
**id**. The `name` is a label for people and for what a manifest writes; it is resolved to an id
once, when a claim is registered. So:
- **A rename is one `UPDATE seats set name = … where id = …`.** Nothing is recompiled, nothing is
re-registered, nothing is refused, nothing freezes. Held records and claims already point at the
id, so they follow the rename for free. The build machine is not involved, because the build
machine validates a claim against the set it reads from the mesh, not one baked into its image.
- **Adding or removing a seat is an `INSERT`/`DELETE`** (within the closed-set discipline: a change
to the set is still a decision with a record — the record is now a row's `decided` column and an
ADR, not a line of Go). No controller release is needed to change the roster of roles.
- **Production code stops hardcoding names.** `const gitSeat = "git"` becomes a lookup of the seat
that delivers the `git` provision (or a well-known id), so renaming its label cannot break the
code that finds a repository's forge.
**What does not change** (0110 and 0121 still hold): a seat is still a module assignment from a
closed set; there is still one holder per scope; a delivering seat is still the single answer for
its provision; system seats are still `mesh-*`/`node-*` and a module may still define its own. Only
their *storage and reference* change — from a compiled slice keyed by name to a table keyed by id.
**A manifest still claims by name, and that is fine.** A manifest is written by a person and names
the seat in words; the mesh resolves the name to an id at registration and stores the id. If a
seat's name changes, manifests written against the old name are updated in the catalogue like any
other edit (and the mesh can keep the old name as an alias row during a transition so nothing breaks
in the window) — but the *control plane* never has to change or redeploy for it, which is the whole
point. The heavy, mesh-wide, freeze-prone half of a rename disappears; only the ordinary catalogue
edit remains.
## Consequences
- **A rename, and a set change, become operations, not releases.** The pain this session paid —
three freezes, a builder deadlock, hand-resolved manifests — is designed out. The seat migrations
still outstanding (the delivering registry seats, and the private network's scope change) should
wait for this: done as data, each is a write, not a coupled multi-repo deploy.
- **The controller gains a small table and a seed migration**, and its seat lookups change from
slice scans to id-keyed reads. `SeatNamed`, `SeatDelivering`, `claimProblems` read the table.
- **The build machine reads the set from the mesh** (it already talks to the control plane), rather
than embedding it — which removes the controller/builder seat coupling that made every breaking
seat change a two-sided deadlock (see [to-be 30](../03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md)).
- **The closed set is still closed.** Data being editable is not the set being open: changing it is
still a decision, still recorded. What changes is that recording it no longer means shipping a
binary.
- **This is a real refactor**, touching the store schema, the seat lookups, claim registration
(name→id resolution), and the held-seat records. It is worth its own build; until it lands, the
current compiled set stands and further renames are held rather than forced through the heavy path.
- **Config on a seat is still the module's** (the question that surfaced this): a seat row carries
the seat's own metadata (scope, delivers, protocol), not a module's configuration — that stays in
the holding module's manifest ([ADR 0046](0046-a-module-configuration-is-its-assignments-not-its-manifest.md)). Making
seats data does not make them a config store.
## References
- [ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md),
[ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md) — the seat
rules this keeps, whose *storage* it changes
- [to-be 30](../03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md) — the controller/builder
seat coupling and the breaking-change freeze this removes for seat changes
- mesh-controller `internal/catalogue/seats.go` (the compiled slice this replaces),
`cmd/mesh-controller/source.go` (`const gitSeat`, the hardcoded name this removes)
@@ -0,0 +1,133 @@
---
topic: the mesh
status: superseded
superseded-by: 02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md
date: 2026-09-26
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0106-the-bus-is-nats.md
---
# 125. The bus is the only broker
## Context
[ADR 0106](0106-the-bus-is-nats.md) moved the mesh's bus to NATS and kept the AMQP broker "as a
module with one purpose — the predecessor's clients", retiring with the last of them.
[Design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md) repeats that: a compatibility module with
a single purpose and a retirement condition.
**It is not single-purpose, and was not when that was written.** Two modules of the *new* mesh
declare `requires: ["amqp"]` and are answered by the broker module's own provisioner:
- `amqp-ping`, whose source says it "exists to PROVE the grant end to end: the mesh gave it a
scoped login and a vhost of that name on the lavinmq provider";
- `amqp-email-forwarder`, which uses it for work.
What that provisioner answers is **not the mesh's bus**. Its own comment draws the line: a
consumer gets "its own message broker, isolated from every other consumer's by the vhost
boundary… a broker of its own, not a shared account on the mesh's control-plane broker" —
vhost-per-login, "the exact analog of postgres's database-per-login."
So two different things wear the word *broker*: the mesh's nervous system, and a private message
broker handed to a module as a resource, the way a database is. The first is being replaced. The
second was never examined, and on the retirement condition ADR 0106 sets, it disappears with no
successor and nothing notices — a module of the new mesh left requiring something no provider
answers.
The operator's direction, asked at the point this surfaced: **NATS is the heart of the
application** — not a component it contains, and not a thing to reproduce the predecessor's
shapes on.
## Considered Options
1. **Carry the private broker forward onto NATS** — each requiring module gets its own NATS
account, provisioned like a database. Rejected on three counts. It reproduces the
predecessor's shape on the new bus, which is the thing this whole move exists to stop. It
gives the mesh two messaging models, so "how does a module send a message" has two answers
depending on a manifest line. And NATS accounts isolate subject spaces *entirely*: a module
inside its own account cannot reach the mesh's bus at all, so it would hold two connections
and two identities to do one job.
2. **Keep the compatibility broker indefinitely** for the mesh's own modules. Rejected: its
retirement condition is the point of it. A module of the new mesh depending on the retired one
keeps the predecessor alive permanently, which is the opposite of a compatibility module.
3. **One bus. A module's messaging is subjects on it, scoped by what it declares.** Adopted.
## Decision
**The bus is the only broker.** NATS is the mesh's one messaging system, and every module's
messaging is subjects on that bus under its own account, scoped by its `emits` and `consumes`
([ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)). There is no second
broker, and none is handed to a module as a resource.
**The `amqp` interface is not carried forward.** It leaves the set of things a module may require
and retires with the compatibility broker rather than gaining a successor.
Concretely, in the controller's seat table: **the `mesh-broker` seat delivers nothing.** It
currently reads `Delivers: "amqp"` — the seat's holder answers a requirement for a broker — and
under this decision it joins `mesh-controller` and `the-catalogue`, the foundation seats that
deliver no provision at all. The bus is not something a module asks for; it is what a module is
reached through.
- `amqp-email-forwarder` moves to the bus like any module: what it emits and consumes, declared,
and the account follows.
- `amqp-ping`'s *purpose* is kept and its mechanism is not. Proving end to end that a module
receives scoped messaging it did not configure itself is worth a probe; it becomes a probe of
the bus, and its assertion changes from "I reached my own vhost" to "I reached exactly my
subjects and was refused the rest."
**A module that wants a queue of its own has one already**: a subject nothing else may publish to
and a durable consumer of its own, both derived from its declaration. What it does not get is a
server of its own.
**The mesh's own streams are the controller's, created at genesis, not provisioned** — and
`EVENTS` is one stream, closing the question [design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md)
§11 left open. The reason is not preference but **bootstrapping**: a provisioner is a module, and
a module needs a bus account before it can run at all. Anything the bus itself is made of must
exist before the first module starts, so it is composed as configuration
([ADR 0106](0106-the-bus-is-nats.md): never through a management API) rather than provisioned by
something that could not yet be running.
## Consequences
- **Design 25 gains the distinction and loses the "single purpose" claim**; its §11 question about
the `EVENTS` stream closes here.
- **Nothing in [design 27](../03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md)'s
model changes** — the four provider kinds, the contract, resolution all stand, and it never
enumerated interfaces, so there is nothing to strike from it. What changes is that messaging
leaves the set of things resolved at all: every module has it by existing.
- **One line of the controller's seat table changes**, and it is the load-bearing one:
`mesh-broker` stops declaring what it delivers. A requirement for `amqp` then resolves to
nothing and is refused at assignment, which is how the two modules below are found rather than
discovered at runtime.
- **Two modules have conversion work**, and it belongs to step 4 of
[ADR 0116](0116-the-bus-is-built-in-five-steps.md), with the flows. Neither blocks step 1.
- **The compatibility broker becomes what ADR 0106 already called it** — single-purpose — once
those two have moved. That record's claim was wrong when written and is made true by this one.
- **What got harder:** a module that genuinely wanted an isolated server — a tenant boundary at
the broker rather than at the subject — no longer has that option, and would have to argue for
it as a new decision. That is the intended cost: one bus is the point.
## How it is checked
- **A module's messaging works with no `requires` line for it.** A lab bed: a module declaring
only `emits` and `consumes` reaches its subjects, and is refused every other — which is
[ADR 0116](0116-the-bus-is-built-in-five-steps.md) step 1's permission bed, already required.
- **Nothing requires `amqp`.** With the seat delivering nothing, a module still declaring it is
refused at resolution — the existing "requirement no provider answers" path, not a new check. A
catalogue test asserts no module declares it once the two have moved.
- **The probe proves the claim it is named for.** `amqp-ping`'s successor fails if a module can
reach a subject outside its declaration, not merely if it cannot reach its own.
- **The compatibility broker's retirement condition can actually be met.** A check that no module
of the mesh — as opposed to a predecessor client — holds a connection to it.
## References
- [ADR 0106](0106-the-bus-is-nats.md) — the bus is NATS; corrected here on what the compatibility
broker serves.
- [ADR 0116](0116-the-bus-is-built-in-five-steps.md) — the steps; the conversions land in step 4.
- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — the scoping that
makes one bus safe.
- [design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md),
[design 27](../03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md) — the two documents
this changes.
@@ -0,0 +1,145 @@
---
topic: the tiers
status: accepted
date: 2026-09-26
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md
---
# 126. A module declares its own seats; the mesh reserves its own
## Context
[ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) closed the set of seats. Its
evidence was strong and still is: nothing could answer *which seats does this mesh have, and who
holds each*. Answering it meant reading every manifest in two repositories and then the
controller's own code, and when that enumeration was done by hand while writing the record, **it
reported eleven claims where there were thirteen.** The fix was a table in the controller, and
adding a seat became a decision.
What that table cannot express is the architecture [ADR 0125](0125-the-bus-is-the-only-broker.md)
opened. With one bus and no private brokers, a module offering a service to other modules offers
it as **a role on the bus**: a set of subjects, exactly one holder, addressed by what it does
rather than by which module or node provides it. A telegram sender, a licensing master, anything
a mesh might want one of. Under a closed table, adding any of those means editing the controller
— so a capability contributed by a module would require a change to the mesh itself, which is the
coupling the module system exists to prevent.
**The two requirements look opposed and are not.** 0110 needs the set *enumerable*. The
architecture needs it *extensible*. Those conflict only if enumerable means *written down in one
place by hand* — which is exactly the property that let the count drift in the first place.
## Considered Options
1. **Keep the closed table, add each new seat by decision.** Rejected. Every capability a module
contributes would need a change to the controller and a record before it could be offered, and
the mesh would carry the names of services it does not itself implement.
2. **Free-form seats, as before 0110.** Rejected for 0110's own reason, unchanged: nothing can
say what a mesh has, and a name invented at a claim site is a name nobody can explain later.
3. **A set that is closed at any moment and derived rather than maintained**, with the mesh's own
seats reserved by name. Adopted. 0110 weighed options 1 and 2 and never considered this one.
## Decision
**A seat may be declared by a module, and the set of seats a mesh has is derived: the mesh's own,
plus those declared by every module it has registered.** The set is still closed — a seat named
nowhere is refused — but it is computed from the catalogue rather than written in the controller.
Everything 0110 decided about what a seat *is* stands untouched: one holder at its scope; a
definition says which seats a module *can* hold and an assignment says which it *does*; holding
one may deliver a provision; a seat makes a role singular, never a module.
**Enumeration is a query, not an inventory.** The catalogue knows every registered manifest, so
"which seats does this mesh have, and who holds each" is answered by asking it. This is a
stronger answer than the table gave, not a weaker one: a derived list cannot drift from reality,
and drift is how the hand-made count came out at eleven of thirteen.
**The mesh's own seats are reserved by prefix.** Every seat the mesh itself defines is named
`mesh-*`, and a module declaring any `mesh-*` name is refused at registration. The prefix *is*
the reservation rule — no list of reserved names to maintain, and no way for the mesh's own
namespace to be colonised by a manifest. This requires renaming the seats that drifted from
[ADR 0079](0079-the-foundation-seats-are-named-after-their-servers.md)'s convention: `the-catalogue`
becomes `mesh-catalog`, `git` becomes `mesh-git`, and the node-scoped `the-build-machine`,
`the-dns-port`, `the-intrusion-prevention`, `the-packet-filter`, `the-private-network`,
`the-resolver-configuration`, `the-showcase` take the same prefix.
The mesh's seats stay the mesh's for a reason that does not apply to a module's: **the mesh's own
code looks them up by name.** The resolver *is* the thing that finds the store. `mesh-store` is
not a convention the controller follows, it is an identifier the controller dereferences.
**A declared seat carries a protocol.** A module declaring a seat says what may be sent to it,
what it emits, and what it serves. The holder must satisfy it; a module may not claim a seat whose
protocol it does not implement. Callers declare that they use the *seat*, never the module, so
replacing the implementation changes nothing for any caller.
**A seat is for a role; an event stays addressed to its emitter.** The two are not
interchangeable and the choice is not stylistic. An event is *this happened to me* — the emitter's
identity is the meaning, which is why the envelope carries source, node and time
([ADR 0042](0042-the-shape-of-an-event-on-the-wire.md)); routing it through a role would erase the
provenance an audit needs. A seat is *this capability, whoever provides it* — where not knowing
the holder is the point. Publish an event when the fact is about you; declare a seat when you are
offering something another module could offer instead.
**Two modules declaring the same seat name is refused at registration**, second one loses.
Registration is the last moment the mesh can still say no, and a seat name meaning two different
protocols is the failure nobody could diagnose afterwards.
## Consequences
- **The controller's seat table stops being the set** and becomes the mesh's own reserved entries.
Resolution reads the catalogue for the rest.
- **Ten seats are renamed.** A rename is a migration, not an edit: existing assignments hold the
old names, so the change carries a mapping and is applied once, and the lab beds that name seats
are updated with it.
- **A `uses` naming an undeclared seat is refused at registration**, which is where 0110's
guarantee lands under this model — the same refusal, at the same moment, from a derived set.
- **Adding a capability stops requiring a decision record.** That is a real loss of governance and
the intended trade: the argument for a seat's existence moves into the module that declares it,
where it is reviewed as part of the manifest. The mesh's own seats keep the old bar.
- **[ADR 0041](0041-events-are-a-relationship.md)'s machinery claim is already stale** for a
different reason, and is corrected in place there under the rule in
[`README.md`](README.md) — a progressive insight: on JetStream a subscription is a durable
consumer, a real object someone must create.
- **What got harder:** a seat's protocol is now a compatibility surface between modules that do
not know each other. Changing one breaks callers already bound to it, and nothing here says how
that is versioned. It is the first thing to answer in the design, and the thing most likely to
hurt later rather than now.
## How it is checked
- **The overview answers, and is right.** A command lists every seat, its scope, its protocol and
its holder, derived from the catalogue — and a test asserts the count against a fixture mesh,
because an enumeration nobody checks is how thirteen became eleven.
- **`mesh-*` is refused to a module.** A registration test: a manifest declaring `mesh-anything`
is refused, naming the prefix as the reason.
- **An undeclared seat is refused.** A registration test on `uses`, and a resolution test that
nothing reaches runtime unresolved.
- **A second declarer loses.** A registration test: two manifests, same seat name, the second
refused and the first untouched.
- **A holder must satisfy the protocol.** A claim whose module does not serve what the seat
declares is refused at assignment, not discovered when a caller times out.
## Progressive insight
> **Progressive insight — 2026-09-26.** *A seat rename is not a data migration.* This record's
> consequences say "a rename is a migration, not an edit: existing assignments hold the old
> names, so the change carries a mapping and is applied once". Implementing it showed there is
> nothing stored to migrate: a seat's holding is **derived at resolution** from the claims in
> manifests (`resolve.go` builds it each time), never written down, so no recorded name is left
> pointing at the old one. What exists is source — the controller's seat table, the manifests
> that claim them, and a manifest that may be registered later from its own repository. So the
> change is an edit plus a **kept** rename table, which tells a manifest written against an old
> name what it became rather than refusing it as unknown.
>
> The decision — that modules declare seats, that the mesh reserves `mesh-*`, and that the ten
> are renamed — is unchanged. Only the shape of the work was wrong.
## References
- [ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) — superseded here; its
requirement is kept and only its mechanism replaced.
- [ADR 0125](0125-the-bus-is-the-only-broker.md) — one bus, which is what makes a role addressable.
- [ADR 0079](0079-the-foundation-seats-are-named-after-their-servers.md) — the naming convention
the reserved prefix restores.
- [ADR 0041](0041-events-are-a-relationship.md) — the event half of the boundary drawn here.
@@ -0,0 +1,104 @@
---
topic: the mesh
status: superseded
superseded-by: 0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
date: 2026-09-26
deciders: jochen
reconstructed: false
supersedes: 02-DECISIONS/0125-the-bus-is-the-only-broker.md
---
# 127. AMQP is a provision, not the bus
## Context
[ADR 0125](0125-the-bus-is-the-only-broker.md) decided that the bus is the only broker, and went
one step further than it had grounds for: it also decided that the `amqp` **interface** — a module
requiring a message broker of its own — "is not carried forward" and "retires with the
compatibility broker rather than gaining a successor", with the two modules declaring it converted
to the bus in step 4.
The operator's correction: **AMQP is deprecated as the mesh's transport, not abolished as a
service.** The broker module keeps running and keeps answering `amqp` requirements. It is no
longer a core part of the mesh — *"it's just a module like mssql now."*
**What 0117 conflated** is two different reasons a module might ask for a broker, which look
identical in a manifest:
1. **To talk to other modules.** Wrong under one bus, and the thing 0117 was right to refuse: a
private broker used as inter-module transport is a second bus, with every guarantee crossing a
seam and no scoping the mesh can see.
2. **Because it genuinely needs an AMQP broker**, the way something needs a database — a queue for
its own internals, or interop with software that speaks AMQP and nothing else. That is a
backing service, and the mesh has a word for backing services already.
0117 saw the first and legislated against both. The second is ordinary, and forbidding it would
make the mesh unable to run a large class of perfectly normal software while claiming that as
architecture.
## Considered Options
1. **Keep 0117 as written** — retire the interface, convert the two modules. Rejected by the
operator, and wrongly reasoned besides: it treats "needs an AMQP broker" as always a mistake.
2. **Keep the broker as the predecessor's compatibility module**, as ADR 0106 framed it, with a
retirement condition. Rejected: it is not single-purpose and its clients are not only the
predecessor's, so the retirement condition describes a day that will not come.
3. **The broker is an ordinary provider module of an ordinary provision.** Adopted.
## Decision
**The mesh's bus is NATS and only NATS.** Everything 0117 decided about *the bus* stands: one bus,
a module's messaging is subjects on it scoped by what it declares, no module is handed a bus of
its own, and the `mesh-broker` seat is the NATS server's.
**`amqp` remains a provision a module may require**, answered by the broker module the way
`postgres-database` is answered by the store module or a database is answered by mssql. It is not
deprecated as an interface; the software behind it is simply no longer the mesh's nervous system.
**The broker module stops being foundation.** It claims no seat — `mesh-broker` is the NATS
server's — it is not raised at genesis, nothing in the mesh requires it, and a mesh that never
installs it is a complete mesh. It is installed when something wants it, like any other provider.
**The rule that survives, stated so it can be applied:** *inter-module communication goes over the
bus.* A module may hold a broker, a database or a cache as a backing service; it may not use one
as a channel to another module. The line is not which software is involved, it is whether a second
module is on the other end.
**Neither `amqp-ping` nor `amqp-email-forwarder` needs converting.** 0117 put that work in step 4;
it is removed. They require a backing service and a provider answers.
## Consequences
- **The "compatibility broker" framing is wrong and goes.** There is no `lavinmq-compat`, no
single purpose and no retirement condition. Design 25 §5 is corrected.
- **[ADR 0106](0106-the-bus-is-nats.md)'s progressive insight was itself wrong** and is corrected
by a second one there. It said 0117 would make 0106's "one purpose — the predecessor's clients"
sentence true by moving the mesh's modules off. Nothing moves off; the sentence is simply not
what the broker is.
- **The seat change stands**, for a better reason than 0117 gave: not because a broker cannot be
provisioned, but because *this* broker is not the mesh's bus. The broker module drops its
`mesh-broker` claim and the `nats` module takes it.
- **Step 4 loses two conversions**; step 1 and the WBS are otherwise unaffected.
- **What got harder:** the rule is now a judgement rather than a prohibition. "Is this a backing
service or a channel to another module?" has to be asked in review, where 0117 could have
answered it with a parser. That is the honest cost of allowing the legitimate case.
## How it is checked
- **A module's own messaging needs no `requires`.** The check from 0117, unchanged: a module
declaring only `emits` and `consumes` reaches its subjects and is refused every other.
- **The broker holds no seat.** A manifest test: the broker module claims nothing, and a mesh
raised without it is complete — genesis names it nowhere.
- **`amqp` resolves like any provision.** A resolution test: a module requiring it is answered by
the provider, refused when none is assigned, and neither case touches the bus.
- **What cannot be checked mechanically**, and is said rather than implied: that a module holding
a broker is not using it to reach another module. Review, not a parser.
## References
- [ADR 0125](0125-the-bus-is-the-only-broker.md) — superseded; its ruling on the bus is kept
whole and only its ruling on the interface is reversed.
- [ADR 0106](0106-the-bus-is-nats.md) — the bus is NATS; its compatibility-broker framing is
corrected here.
- [ADR 0126](0126-a-module-declares-its-own-seats.md) — seats, including the one the NATS server
now holds alone.
@@ -0,0 +1,123 @@
---
topic: the mesh
status: accepted
date: 2026-09-26
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
---
# 128. The mesh bus is required, not ambient
> **Pointer repointed, 2026-09-27.** This record was written extending
> [ADR 0127](0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)), which
> [ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md) has since superseded — AMQP is
> not a provision at all. Nothing decided here changes; the frontmatter now rests on the live record,
> and the citations below are read with that in mind.
## Context
[Design 29](../03-DESIGN/01-to-be/32-what-a-module-declares.md) opened by saying the bus is
*ambient*: "No module requires it, the way no module requires a filesystem. Every module gets a
connection and an identity whether it asks or not."
**Two counts say that is wrong.** Of the 72 modules in the catalogue, **49 declare an own-secret
named `broker` and 23 do not.** So the bus is not universal — nearly a third of the catalogue
never speaks to it — and an ambient connection would mint an account, a password and a permission
set for every one of those 23, each a credential nothing uses and everything must rotate.
And the 49 that do take one **each hand-write the path it lands at**
(`own-secrets: { broker: "/var/lib/<module>/broker" }`). That is a special case doing badly what
provisioning already does well: a consumer names where a credential lands, the mesh seals it
there, and rotation and removal follow the same path as every other credential.
**The argument that made the bus ambient was narrower than it looked.**
[ADR 0125](0125-the-bus-is-the-only-broker.md) reasoned that bus accounts cannot be provisioned
because a provisioner is itself a module that needs an account before it can run. That is true of
a **provisioner process**, and it is not true of a provision: the mesh's bus accounts are composed
by the *controller*, into configuration, and the controller is not waiting on a bus account to
exist. The circularity is real for one mechanism and absent for the other, and the earlier record
applied it to both.
## Considered Options
1. **Keep the bus ambient.** Rejected on the counts above: it over-grants to 23 modules and keeps
a hand-written path in 49.
2. **Derive the requirement** from whether a module declares any `emits`, `consumes`, `serves` or
`uses`. Rejected: it is the ambient model with extra inference. A reader of a manifest still
cannot see that the module holds a bus credential, and the rule would have to be re-derived
every time the set of bus-facing declarations grew.
3. **The mesh bus is a provision a module requires**, delivered by the seat that holds it.
Adopted.
## Decision
**A module that speaks to the mesh requires `mesh-bus`, and receives what it needs to connect.**
The contract is an address, a credential sealed to the module, and the trust to verify the
server. It lands where the module's manifest says, like any provision. A module that does not
require it gets no account, no password and no permissions — and 23 modules in the catalogue
should get none.
**The `mesh-broker` seat delivers `mesh-bus`.** Its holder is the mesh's own bus, and what
holding it delivers is the connection to that bus — which is what a seat delivering a provision
has always meant ([design 26](../03-DESIGN/01-to-be/26-the-seats.md)).
**The requirement delivers the connection; the declarations shape the authority.** They are two
different things and both stay explicit. `requires: mesh-bus` says *this module talks to the
mesh*; `emits`, `consumes`, `serves`, `uses` and a declared seat say *what it may say and hear*,
and the permission set is derived from those and nothing else
([ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md)). Requiring the bus
grants no subject; declaring a subject without requiring the bus is refused at registration as
incoherent.
**The `mesh-bus` provision is answered by the controller, not by a provisioner.** This is the
surviving kernel of ADR 0125's bootstrap argument, narrowed to what it actually supports: the
bus's accounts are configuration the controller composes and the server reloads
([ADR 0106](0106-the-bus-is-nats.md) — never through a management API), so there is no provisioner
process in the path and nothing waiting on a bus account to create bus accounts. It is a provision
whose provider is the mesh itself.
**A module may also provide a NATS server of its own, and that is a different interface.** Exactly
as the AMQP broker provides `amqp` ([ADR 0127](0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md))), a module
may run its own NATS and offer it as a backing service. That interface is **`nats`**; the mesh's
own bus is **`mesh-bus`**; the two are never the same name, because a manifest that said `nats`
could mean either and the difference is the whole architecture. The rule from 0119 decides which
is legitimate: a private bus is a backing service, never a channel to another module.
## Consequences
- **Design 29's opening is reversed.** The bus is not ambient; it is required, and the document's
first paragraph says the opposite of this.
- **The seat's `Delivers` is `mesh-bus`** — corrected twice in one day, which is worth recording
rather than tidying: it read `amqp`, which was the old broker's interface; ADR 0125 emptied it,
on the reasoning that a bus cannot be provisioned; and it is neither. The seat delivers the
mesh's bus.
- **`own-secrets: { broker: ... }` is retired** in favour of the provision's own delivery, across
49 manifests. That is a mechanical change, and it belongs with the conversions in step 4 rather
than step 1.
- **23 modules lose a credential they never used.** Not a regression — an over-grant removed, and
the smallest honest statement of what this buys.
- **What got harder:** one more line in most manifests. The trade is that the line is true, and
its absence is also true.
## How it is checked
- **A module with no `requires: mesh-bus` has no account.** A composition test: the derived user
list contains exactly the modules that require it, and the 23 that do not appear nowhere in it.
- **Declaring a subject without requiring the bus is refused.** A registration test on a manifest
with `emits` and no requirement, naming the contradiction.
- **Requiring the bus grants no subject on its own.** A composition test: a module that requires
`mesh-bus` and declares nothing else gets a connection and an empty permission set.
- **`nats` and `mesh-bus` are distinct interfaces.** A resolution test: a module requiring `nats`
is answered by a module providing it, never by the seat holder, and vice versa.
## References
- [ADR 0125](0125-the-bus-is-the-only-broker.md) — superseded by 0119; its bootstrap argument is
narrowed here to the case it supports.
- [ADR 0127](0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)) — a broker as a backing service; this
applies the same shape to the mesh's own bus and separates the two names.
- [ADR 0043](0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) — authority from
declarations, which this leaves untouched.
- [design 26](../03-DESIGN/01-to-be/26-the-seats.md) — a seat delivering a provision.
@@ -0,0 +1,102 @@
---
topic: the mesh
status: accepted
date: 2026-09-27
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0122-a-seat-is-data-a-rename-is-a-database-update.md
---
# 129. A seat carries the protocol of its role
## Context
[ADR 0126](0126-a-module-declares-its-own-seats.md) let a module declare a seat with its protocol:
what work the role accepts, what it emits, what it serves. A module's own seats work that way today.
**The mesh's own seats — the `mesh-*` set — carry no protocol at all**, only a name, a scope and the
provision they deliver. They say who does a job and nothing about what may be said to them or by
them.
That gap surfaced three times in one day, each time as a different-looking problem.
**A build machine.** On the bus the mesh runs on today a builder has its own account kind, created by
its own command, with permissions written by hand: read the build queue, write to two exchanges. One
publish to a shared exchange reached all three audiences a finished build has — whoever asked, the
controller that records it, and the catalogue that places it in the module graph. On a bus where
permissions are per subject those are three separate grants, and nothing derives them, because a
builder is not a module and holds a seat that promises nothing.
**An event about a role rather than about a module.** The module holding the artifact-store seat
declared an event named after a *different* module
([issue 127](../04-ISSUES/127-a-module-event-derives-a-subject-nothing-publishes/00-report.md)). The
bus refuses that, because a namespace belongs to who it is named for. The event is genuinely about the
role — "the artifact store accepted an image" — and a consumer written against whichever module holds
that role today breaks when the holder changes. There was nowhere else to put it.
**A catalogue catching up.** The controller answers a request for builds it may have missed by
re-publishing them under its own name, which no consumer of the builder's subject hears. Publishing
them under the builder's name would be the controller signing an event as another module. Answering
into the asker's inbox needs a grant over every inbox in the mesh, which
[design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §4 refuses.
Three symptoms, one cause: **the mesh has roles it cannot describe.**
## Decision
**A seat carries the protocol of its role, whether the seat is a module's or the mesh's own.** The
`mesh-*` set gains the same three fields a declared seat has — what it accepts, what it emits, what it
serves — and the holder's authority, its work queue and its consumers are derived from them by the
machinery that already does this for a module's seats.
**Builds become work submitted to a role.** The build machine seat accepts a build and emits an
outcome. The dedicated `mesh.build.*` branch and the stream behind it retire: a work queue shared by
several build machines is exactly what a seat's `accepts` already is, and keeping a second mechanism
for it means two things to reason about and two places for a permission to be wrong.
**One publish still reaches three audiences, and now the mesh derived the subject.** A build's outcome
is the seat's own event. Whoever asked matches it by the id their request carried; the controller
records it; the catalogue places it. That is the fan-out the shared exchange gave for free, expressed
as a subject rather than as a topology, and it means no holder needs permission to publish into
anybody's inbox.
## Alternatives considered
**A dedicated principal kind for a builder**, mirroring the account the old bus issues it. Smaller: one
addition to the composer, no change to seats, and it matches how a builder is treated today. Not taken
because it answers one of the three symptoms and leaves the other two, and because "the builder is
special" is a claim nobody could justify from the design — a build machine is a role the mesh has, and
the mesh has a word for a role.
**Leaving the outcome as a reply to the asker's inbox.** Rejected on authority: a holder able to answer
any asker needs a grant across the whole inbox space, which is the one grant design 25 §4 refuses by
name. The seat's event costs the asker a filter and costs the mesh nothing.
## Reconciled with 0122, which landed in parallel
*Added 2026-09-27, on merging.* [ADR 0122](0122-a-seat-is-data-a-rename-is-a-database-update.md) moved
the seat set out of compiled code and into a table the controller owns. This record was written against
the slice, and says the `mesh-*` set "gains the same three fields a declared seat has".
**The decision is unaffected and the mechanism is better for it.** What a seat accepts, emits and serves
becomes three columns beside its name and scope, so giving a role a protocol is a write rather than a
rebuild — which is the whole argument of 0122 applied to the thing this record adds. Where this text
says the set gains fields, read: the table gains columns.
## Consequences
**A seat is now the mesh's unit of "a role that talks".** A role that accepts work, announces outcomes
or answers questions says so where it is defined, and everything about permissions, queues and
consumers follows. Nothing hand-writes a grant for a role again.
**The shared library cannot yet publish on a seat, and that is now the blocking gap rather than a
curiosity.** A module holding a seat has the authority and no way to use it; the build machine is
written in Go and reaches the bus directly, so it is unaffected, but the artifact-store event stays
under its module's own name until the library has a surface for this. That is a task, and this record
is what makes it one.
**A second mechanism disappears.** `mesh.build.*`, the BUILDS stream and the builder's hand-written
account all retire. Fewer things, and the ones left are derived.
**The catch-up question is not settled by this**, only made answerable: a seat that serves something
gives the controller a way to be asked, which the mesh did not have. Whether catch-up should be a
question at all remains open.
@@ -0,0 +1,71 @@
---
topic: the mesh
status: accepted
date: 2026-09-27
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
---
# 130. The predecessor is ending, and its broker goes with it
> **Pointer repointed, 2026-09-27.** This record was written extending
> [ADR 0127](0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)), which
> [ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md) has since superseded — AMQP is
> not a provision at all. Nothing decided here changes; the frontmatter now rests on the live record,
> and the citations below are read with that in mind.
## Context
[ADR 0127](0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)) settled that the old broker is an ordinary
provider of the `amqp` provision rather than a compatibility module with an end date. It rejected
giving it a retirement condition, and said why: *"its clients are not only the predecessor's, so the
retirement condition describes a day that will not come."*
**The operator has said that day is coming.** The predecessor is deprecated. Some of it is still
running, and it is not being migrated — it is being left to stop. Its broker may be shut down.
That is a fact about this installation, not a change of mind about what a broker is. It is recorded
because three documents reason from the premise it overturns:
[design 25](../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §5 and §9, and
[design 28](../03-DESIGN/01-to-be/28-building-the-bus.md)'s closing note that the predecessor's world
"does not need to move: its broker is the compatibility module until its last client is gone."
## Decision
**The predecessor's broker retires when nothing requires `amqp`, by being unassigned like any other
provider.** No retirement condition, no end-date machinery, no special case — which is ADR 0127 being
paid off rather than revised. Because that record made the broker an ordinary provider, ending it
needs nothing that does not already exist: a provision with no consumers has its provider unassigned,
and the module system has done that since it existed.
**So step 5.3 has an ending.** "The mesh's own accounts removed from the deprecated broker" was
written as the last thing that could be said, because the broker itself was going to outlive the
question. It now finishes: once the mesh's own traffic has moved and the predecessor's remnants have
stopped, the module is unassigned and the port is free.
**And the transitional doubling has a date.** The build outcome is announced under both the module's
name and the role's on the old bus, so that a catalogue deployed before the rename and one deployed
after both hear it. That exists only while the old bus does, and goes with it.
## Consequences
**The remote access path goes with it, and that is the one practical consequence worth planning
around.** The predecessor's own mesh communicates over that broker — so shutting it down ends the
tooling that reaches this installation's machines remotely. Work on the node after that point is done
from the node. **This matters most for the rollout**, which is the step that would otherwise be driven
from a workstation: it has to be driven locally, or driven before the broker stops.
**What is still running on it stops when it stops.** Some of the predecessor's services are live and
are not being moved. That is the operator's decision and it is recorded here so that nobody later reads
a broker with clients as an accident.
**Nothing in a served request's path is affected.** Modules serve from their own containers; the mesh's
bus carries the mesh's own traffic — declarations, reports, events, tool calls. This was checked rather
than assumed when the question came up, and it is why the operator's position (*"as long as my services
keep running"*) is a bounded risk rather than a gamble.
**One reason to keep the broker survives**: `amqp` remains a provision a module may require, and a
module that genuinely needs an AMQP broker can be given one. What retires is *this* broker's role as
the predecessor's, not the mesh's ability to provide the thing.
@@ -0,0 +1,94 @@
---
topic: the mesh
status: accepted
date: 2026-09-27
deciders: jochen
reconstructed: false
supersedes: 0127-amqp-is-a-provision-not-the-bus.md
---
# 131. Everything on the mesh speaks to the broker seat, and AMQP is not a provision
## Context
[ADR 0127](0127-amqp-is-a-provision-not-the-bus.md) settled the old broker as an ordinary provider
of an ordinary provision, `amqp`, kept for whatever wanted a message broker of its own. The day the
bus moved was the day that framing was tested, and it failed in a way that took the control plane
down for an evening.
Three things came out of the wreckage. **The protocol had leaked into the seat's contract**: for a
module to hold `mesh-broker`, it had to provide what the seat delivers, and what it delivered was
`amqp` — so the module that will carry the bus on NATS could not hold the seat that names the bus,
while the module the mesh was leaving could. **A consumer of `amqp` is not asking for AMQP.** The two
modules requiring it wanted the mesh's messaging — to emit an event, to hear a topic — and named the
wire protocol only because that was the word available. **And AMQP and NATS are not interchangeable
at the wire.** A provision named after a protocol can only ever be answered by that protocol, so once
the bus is NATS an `amqp` provision has one possible provider, and it is the thing being retired.
The operator's position, stated during the outage: modules depend on the broker *seat*, not on a
protocol; AMQP is obsolete as anything the mesh's core knows about; a module that depends on `amqp`
is wrong; and everything should reach the mesh's bus and be able to emit events and consume topics
through it.
## Decision
**A module that needs messaging uses the mesh's bus, and the mesh's bus is whatever holds
`mesh-broker`.** Emitting an event and consuming a topic go through the sdk, which is handed the
bus by the mesh with the module's own credential. No manifest names a wire protocol to get it.
**`amqp` is neither a provision nor a requirement.** Registration refuses a manifest that provides
it or requires it. The `mesh-broker` seat delivers `mesh-bus`, and its holder is the module that
provides `mesh-bus` — today the nats module, and only it.
**The old broker's module and the two modules that required it leave the catalogue.** They are
removed, not converted: one was a proof that a grant worked end to end, the other forwards mail off a
queue, and both are re-done against the bus if wanted, as new modules under this record.
**The controller's AMQP transport is deleted once every node reports on the new bus**, and the
switch that selects a transport goes with it — one bus, so nothing to select.
The predecessor's own broker is outside the mesh and not this record's concern
([ADR 0130](0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md)): what the predecessor's
tooling loses when it stops is accepted there.
## Options considered
1. **Keep 0127: AMQP stays an ordinary provision with the old broker as its provider.** Rejected. It
is what put the protocol into the seat's contract, it is why the seat could be left with no valid
holder mid-change, and it keeps two transports in the control plane indefinitely for the benefit of
two modules that did not want AMQP in the first place.
2. **Bridge it: the old broker's module also provides `mesh-bus`, so both can hold the seat during the
change.** Rejected. It makes the retiring broker a legitimate mesh bus for exactly as long as
nobody removes the line, which in practice is forever, and it leaves `amqp` as a thing the core
still knows the name of.
3. **The seat is the dependency; the protocol is nobody's business but the holder's.** Adopted.
## Consequences
- **The change of holder is a handover, and it needs a command.** Nothing today moves a seat from
one assignment to another as one act, and a seat the control plane dereferences cannot be empty
in between — that emptiness is the outage this record comes from. The command takes a seat and the
assignment taking it over. Designed and built before the cutover, under
[28 — Building the bus](../03-DESIGN/01-to-be/28-building-the-bus.md).
- **The seat's row moves to `mesh-bus` before the new holder registers, and that is safe.** The
control plane composes its own bus address through the seat *by name*
(`${seat:mesh-broker:…}`), and the overview derives holders by name; only registration and the
provision-to-seat resolution read what a seat delivers. So the row can change under the current
holder without unseating it, the new holder can then register its claim, and the handover happens
when both are running. Verified in the code during the outage, not assumed.
- **Registration gains two refusals**: a manifest providing `amqp`, and one requiring it.
- **The `rollout check` stops saying the old broker stays.** It said so under 0127; it now lists
unassigning it as the last step of the move.
- **What got harder**: a third party that genuinely wants an AMQP broker on a mesh node runs one as
any application module, with no provision and no seat, and nothing on the mesh routes to it. That
is the cost of the mesh not knowing the word.
## How this is checked
| Rule | Checked by |
|---|---|
| No manifest provides or requires `amqp` | a registration test refusing each, naming this record; and a whole-catalogue test asserting no registered manifest names it |
| `mesh-broker` delivers `mesh-bus`, and only a `mesh-bus` provider may hold it | the existing registration test for a delivering seat, with the row's value read from the store (mesh-controller#89) |
| The seat's row can change without unseating the holder | a test composing the control plane's own address and the overview under a row that the current holder does not satisfy |
| The rollout does not leave the old broker running | `rollout check` output, asserted in its test |
| The AMQP transport is gone | the package does not compile with it referenced; the switch variable is refused as unknown at start |
@@ -0,0 +1,150 @@
---
topic: the mesh
status: accepted
date: 2026-09-28
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md
---
# 132. A seat carries the tools its holder must serve
## Context
[ADR 0129](0129-a-seat-carries-the-protocol-of-its-role.md) gave a seat the protocol of its role in
three parts: the work it accepts, the events it emits, and the verbs it **serves** — request and
reply, awaited. The bus already derives authority from all three: a holder subscribes
`mesh.seat.<seat>.tool.<verb>`, and a module that uses the seat may publish it and nothing else.
**The serving third has never been used.** The mesh defines 14 seats, 8 mesh-scoped and 6
node-scoped. Exactly one carries a protocol at all — the build machine, which accepts `build` and
emits `built`. Not one seat declares a single verb it serves. The mechanism is built, enforced, and
empty.
Meanwhile every tool on the mesh is addressed to a module. Of 72 modules in the catalogue, 45 serve
tools, about 203 of them, each on `mesh.mod.<module>.tool.<name>`. So a caller binds to the module
that happens to hold a role rather than to the role, and replacing that module breaks every caller —
which is the thing seats exist to prevent everywhere else.
**Nothing can say what tools exist.** Measured on 2026-09-28, with the bus carrying the whole mesh: a
workstation client holding an operator credential connected, the bus accepted the account, and
`mesh call gitea.gitea_list_repos` answered with real repositories. The same client's `mesh tools`
found nothing, because it asks `mesh-catalog.catalog_tools` and no module serves that: the catalogue
serves `catalog_modules`, `catalog_module`, `catalog_provides`, `catalog_dependents` and
`catalog_stale`. An agent can therefore call any tool it already knows the name of and discover none.
MCP's `tools/list` is that same question, so the MCP surface is a working transport over an empty
catalogue.
**And there is nowhere for a tool's definition to live.** A manifest has a `tools` field: 0 of the 45
modules that serve tools fill it. That is not neglect, it is the arrangement failing — the field was
the bus grant's source for what a module may subscribe, and because nothing filled it every module
that served a tool was refused its own subscription on the new bus, live, until the grant was changed
to the module's own namespace. Today a tool's name, description and argument schema exist only in the
module's code.
Two facts about the machinery matter for what follows. A seat's protocol is not in the store: the seat
rows lack the ADR 0129 columns, so the protocol comes from compiled defaults and is merged in when a
row is read. And `seatSubject` is flat — `mesh.seat.<seat>.<kind>.<verb>` with no node in it — so a
node-scoped seat's tool call would reach every node's holder at once, and the holders' queue group
would hand it to whichever answered first.
## Decision
**A seat's protocol carries its tools in full**: the verb, what it does, and the schema of its
arguments and of its answer. The seat is the definition of the role's interface; the holder is an
implementation of it.
**Serving the seat's tools is a condition of holding the seat.** A module that does not serve every
verb the seat declares may not occupy it. This is checked where the other conditions of holding are
checked — registration and handover — and refused by naming the verbs that are missing.
**A role's tools are addressed to the role.** `mesh.seat.<seat>.tool.<verb>` mesh-wide. A node-scoped
seat carries the node in the address, because one subject reaching six machines' holders is not an
address, and the queue group that made it look like one would silently pick a winner.
**A module keeps its own tools, and both exist.** `gitea_list_repos` stays, because gitea can run
without holding the `git` seat — a second forge, an instance kept for one purpose. The module's name
answers *this gitea*; the seat's verb answers *whoever is the forge*. Which of the two a caller wants
is a decision in the running session, not one the mesh makes for it.
**What answers "what tools exist" follows where the definition lives.** A seat's tools are read from
the mesh's own records. A module's own tools are answered by the module, from the code that defines
them. Discovery is therefore a read for the durable half and a question to the running mesh for the
free half.
**A seat's tools are an interface, and change like one.** Additive within a version; a change that
would break a caller takes the version token the subject already has room for (design 29 §8), and the
two run side by side until nothing is bound to the old one.
**The mesh's own verbs are the `mesh-controller` seat's tools.** `status`, `push`, `build`, `assign`
and the rest are a role's interface, not a container's, and the audit point [ADR 0095](0095-the-control-plane-is-the-way-to-ask-a-module.md)
asks for is the seat's holder.
## Options considered
1. **The manifest declares each module's tools.** Rejected. The list is then written twice — in the
manifest and in the code — and a schema in a manifest goes stale silently, which is the worst kind
of wrong for something an agent reads to decide what to call. It is also the arrangement that has
already failed once: the field exists, 0 of 45 modules fill it, and the grant that depended on it
refused every tool subscription on the mesh.
2. **Every runtime answers an introspection call, and something aggregates them.** Rejected as the
shape for a role's tools, kept for a module's own. An aggregator needs permission to publish into
every module's namespace, which is a widening the mesh otherwise gives only to the control plane;
and the answer is only as available as the modules are, so a mesh whose catalogue cannot say what a
role answers while its holder is down cannot plan against it.
3. **The control plane answers everything.** Rejected. It puts a tool surface on the control plane for
tools it does not implement, and makes discovery depend on the one component that must stay
answerable while it is itself being replaced. The mesh's own verbs are its to answer, and it answers
them as the holder of a seat.
4. **Seats only; no module tools.** Rejected. Most modules hold no seat, and inventing a seat per
module to give its tools a home would dilute what a seat is: one holder of a role the mesh needs
exactly one of.
## Consequences
**One capability can have two names, deliberately.** A forge that holds the `git` seat answers both
`mesh.seat.git.tool.list_repos` and `mesh.mod.gitea.tool.gitea_list_repos`. This is the one place the
mesh accepts two names for one thing, because they are answers to different questions and the second
one survives the module not holding the seat. The glossary rule stands everywhere else.
**A seat becomes a contract to implement.** Adding a verb to a seat is a change every holder must
make, and a claim that was valid becomes invalid until it does. That is the point, and it is also the
reason a seat's tools should be few and durable while a module's own stay free.
**Three prerequisites, none of them in place.** The seat's protocol must be in the store rather than in
compiled defaults, or discovery reads a binary rather than the mesh. The protocol must become richer
than a list of verbs, because a verb without a schema is not something an agent can call. And a
node-scoped seat needs the node in its subject before any of its tools can exist.
**Discovery becomes cheap for the half that matters.** What roles the mesh has and what each answers is
a query, with no fan-out and nothing to be up. An agent's authority can then be role-shaped — *the
forge's tools* — rather than a list of module-specific names that changes when a module is replaced.
**The MCP surface belongs inside the mesh.** Once the tools are the mesh's own records, the thing that
serves them to an agent is a module the mesh assigns to the machine where the agent sits, with a
credential the mesh minted and authority derived from what it may call — not a program started by hand
with a credential printed to a terminal.
## How this is checked
- **Holding is refused without the verbs.** The condition sits with the other conditions of holding a
seat, so registration and a handover both refuse a module that does not serve what the seat declares,
and the refusal names the missing verbs. A test per condition, as the other seat conditions have.
- **The grant is derived from the seat, and already is.** A holder's subscription and a user's publish
come from the seat's protocol, so a verb nobody declared is a subject nobody may use, and a verb the
seat declares reaches exactly its holder. The golden composition of the bus's user list is the test
that keeps it honest.
- **Discovery is a read, and is tested as one.** What the mesh answers for a seat's tools equals what
the seat's records declare — no call to a module in the path, so the test needs no running module.
- **A node-scoped seat's subject carries its node**, checked by the same test that checks the subject
table: two nodes holding one node-scoped seat derive two addresses.
## References
- [ADR 0129](0129-a-seat-carries-the-protocol-of-its-role.md) — the protocol this widens
- [ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md) — the role, its work and its events
- [ADR 0095](0095-the-control-plane-is-the-way-to-ask-a-module.md) — a tool call passes one process where an audit belongs
- [ADR 0126](0126-a-module-declares-its-own-seats.md) — an event is addressed to its emitter, for the same reason a role's verb is addressed to its role
- [`03-DESIGN/01-to-be/26-the-seats.md`](../03-DESIGN/01-to-be/26-the-seats.md) — how a seat is held and handed over
- mesh-controller #116, #117, #118 — the grants as they now stand: a module serves its own namespace, the control plane may ask any tool
- Measured 2026-09-28 on the live mesh: an operator credential calling a module's tool over the bus answers; `tools/list` finds nothing
@@ -0,0 +1,159 @@
---
topic: what runs on it
status: superseded
date: 2026-09-28
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md
superseded-by: 02-DECISIONS/0135-a-module-version-prepares-its-state-before-it-runs.md
---
# 133. A module owns its migrations, and the mesh owns when they run
## Context
On 2026-09-28 the mesh replaced its own control plane, through its own upgrade path, with a build
carrying a migration. Nothing applied it. For the next three quarters of an hour every build the mesh
made was refused by the store with one line — *column "built_contexts" does not exist* — which reached
only whoever happened to be waiting on that build's reply. The images were built and published, so the
registry filled with artifacts the mesh has no record of, and the overview went on reporting that every
module was current ([issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md)).
The schema had been created once, at genesis, by an action in the foundation bundle. Nothing ran it
again, through many updates of the control plane since.
**The mechanism to do this right already existed and one module used it wrong.**
[ADR 0052](0052-a-step-that-runs-once-before-a-container.md) makes a run-once container a step the host
runs to completion before whatever the declaration places after it, and names migrating a schema as the
case it exists for. Three facts about how it is used today:
- The control plane's manifest had no step at all. The immediate fix was to write one by hand, and that
hand-written step repeats three environment variables and three volume mounts from the server
resource it precedes — six chances to drift from the thing it prepares.
- Two other modules hand-write the same shape for the same reason: gitea's admin bootstrap and
mosquitto's dynsec seed, each repeating its sibling's image, environment and mounts. One of them
ends in `|| true`, which is a lock implemented as a shrug.
- The catalogue module takes the other road: it migrates its own schema in its own code when it starts.
That failure mode is a crash loop rather than a stop — the catalogue restarted 338 times this
morning on an unrelated start-time failure, and nothing anywhere said the mesh's graph had a gap.
**What the mesh already has, and what HAL needed stages for.** Ordering a provider before its consumer
is `providersFirst`, which topologically orders a node's modules. Ordering within a module is
declaration order, and a run-once container gates everything after it. Remembering that a step has
already run is the digest of its declaration, recorded only after it exits 0
([ADR 0018](0018-a-picture-is-read-from-what-runs.md)) — and the image is part of that digest, so a new
build re-runs it. Three of the four things a stage system provides are therefore already here. The
fourth — that a module has a schema at all — is the only thing missing.
**Nothing in the catalogue ships a migrations directory.** Of 72 modules, none has one; the modules that
migrate do it in their own code. So this is not a decision about where SQL files live. It is a decision
about who runs them and when.
**Two facts bound what is safely expressible.** A node converges toward its own declaration without
waiting on any other node. And of the five modules that run on more than one machine today — dnsmasq,
fail2ban, networking, networkmanager, sshd — not one wants a store; every module with a database is on
exactly one machine.
## Decision
**A container may declare steps to run before it.** The same container, run to completion, with
different arguments, in order, before it starts. The mesh derives the run-once resources from that
declaration, so the image, the environment, the volumes, the network and the credentials come from the
one place they are already described and cannot drift from it.
**A module's migrations are the first user of this, and the module owns them entirely.** The SQL, the
order, the idempotence, the lock, and which dialect it speaks. The mesh never learns that postgres and
mssql differ, because it runs the module's own image with the module's own arguments against the
module's own binding and requires exit 0. A module needing both stores runs one step that does both.
**The mesh owns the moment, and the gate is the guarantee.** Whether a version may serve when its
schema is not there yet is a deployment question, and the mesh is the only thing that can answer it,
because the mesh is what starts the container. A step that fails stops the container it precedes, so
a failed migration is a version that does not serve rather than a version serving against a store it
does not match.
**Per node, and there is no level.** The step runs wherever the module runs. A step that ran "once,
somewhere" would leave every other machine with no gate at all, and additive migrations protect old
code against a new schema, never new code against an old one. The cost is an obligation a migration
runner already carries: a version table and a lock.
**"Once, mesh-wide" is what holding a seat means.** A step that is not idempotent — seeding an
account, sending a notice, taking a backup — belongs to a module that holds a seat, where the mesh
already guarantees one holder, on record, handed over deliberately. That is the answer to the level
question rather than a field that has to invent an election and keep it somewhere.
**Migrations are forward-only and additive.** The step runs before the *new* container starts, so the
old one is still running against the new schema for the length of the apply.
**Declared, never inferred.** The control plane cannot see inside an image, so a module that ships
migrations and declares no step is not refusable at registration; it breaks on its first upgrade. This
record says so rather than implying a check that cannot exist.
## Options considered
1. **Each module migrates itself when it starts** — what the catalogue does today. Rejected: it turns a
schema failure into a crash loop instead of a stop, it is invisible in the declaration so nothing can
say the module even has a schema, and two machines running the module both migrate at start with
nothing sequencing them.
2. **The mesh applies migrations itself**, with a driver and a version table per store — HAL's shape.
Rejected: the mesh would have to know one store type from another, hold another module's store
credentials, and reach a machine to use them, which [ADR 0005](0005-the-node-host.md) forbids. It is
also the reason that shape needs levels: something central has to decide where the once happens.
3. **A hook lifecycle** — pre-build, post-build, pre-deploy, post-deploy. Rejected: there is no deploy
event here to hook. A declaration is a desired state applied in order and reconciled forever, so
"pre-deploy" is exactly "a step before this container", pre- and post-build are what a Dockerfile and
the artifact list already are, and "post-deploy" has no moment to name.
4. **A hook level** — once per module, or once per module-node assignment. Rejected as a field, kept as
a property: see the decision. A once-per-module step needs cross-node ordering underneath it to be
safe, and a node converging without waiting on its neighbours is worth losing on purpose rather than
by accident.
5. **Every module hand-writes its own run-once step** — the immediate fix for the control plane.
Rejected as the general answer: it duplicates the resource it precedes, in three places already, and
a hand-written step is one the next module forgets. Forgetting it is the fault this record exists
for.
6. **Record a schema level per module in the store.** Rejected: gating makes the invariant true by
construction, so a level is a second account of the same fact and the first one to go stale.
## Consequences
**Three hand-written steps collapse into one line each**, and the control plane's own migrate step stops
repeating its server's environment and mounts.
**The catalogue's self-migration becomes the exception to remove.** One shape, and the mesh's own
control plane is not an exception to it either.
**A module on two machines with one shared store must lock.** Today none is, so this is an obligation
stated before it is needed rather than discovered by two concurrent migrations.
**There is still no readiness-gated step.** Only an action carries `verify`; a container has no health
notion, so "run this once the service answers" remains unexpressible and seeding through a running
service's API has no home. That is its own decision about a container's readiness, and this record does
not make it.
**Genesis keeps its own action.** At birth there is no control plane to derive anything from, which is
what [ADR 0067](0067-genesis-is-a-pivot.md) already says about that moment.
## How this is checked
- **The composition carries the step.** A test on a node's composed declaration: every container that
declares steps before it is preceded by them, and the derived step's image, environment, volumes and
network equal the container's — so the two cannot drift, which is the failure the hand-written kind
has.
- **A failed step stops what follows.** The host already refuses to go on past a run-once step that did
not exit 0; the test for that is extended to a derived one, so the gate is checked rather than
assumed.
- **The mesh's own schema is covered by the same mechanism as everything else.** The control plane
declares its step in its own manifest, so the case that failed on 2026-09-28 is the case the test
covers.
- **A module claiming a seat for a once-only step is checked where seats are checked** — the conditions
of holding, not a new mechanism.
## References
- [ADR 0052](0052-a-step-that-runs-once-before-a-container.md) — the step this extends
- [ADR 0018](0018-a-picture-is-read-from-what-runs.md) — a digest is the record that something happened
- [ADR 0005](0005-the-node-host.md) — the control plane decides and never touches a machine
- [ADR 0067](0067-genesis-is-a-pivot.md) — why genesis does it differently, once
- [issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md) — the failure that produced this record
- [`03-DESIGN/01-to-be/32-what-a-module-declares.md`](../03-DESIGN/01-to-be/32-what-a-module-declares.md) §6 — the lifecycle this sits in
- Measured 2026-09-28: three hand-written run-once steps repeating their sibling's resource; 0 of 72 modules with a migrations directory; 5 modules on more than one machine, none of them wanting a store
@@ -0,0 +1,126 @@
---
topic: the mesh
status: accepted
date: 2026-09-28
deciders: jochen
reconstructed: false
---
# 134. The mesh says what it applied
## Context
The pipeline is observable on the bus from a merge to an artifact, and modules already plug into it:
the forge emits `pull.merged`, the build machine's seat emits `built`, the catalogue emits `registered`,
`upgraded` and `rebuild-needed`, providers emit `postgres.database.provisioned` and its siblings. Things
consume them today — the catalogue consumes `built`, each provider consumes its own provisioning events,
model-usage consumes `*.usage.*`, the audit logger consumes `**`. Nothing had to be invented for any of
that; subscribing *is* plugging in.
**It goes dark at the moment it touches a machine.** A host applies a declaration and reports to the
control plane on the control branch, which only the control plane may read — correctly, because a report
carries what a machine is and enrolment travels the same way. So nothing on the mesh says *this machine
now runs version Y of module Z*, or that it refused to, or why.
What that cost on 2026-09-28, in one morning:
- A build result the store refused was visible only to whoever was waiting on that build's reply. For
three quarters of an hour the mesh built things and recorded none of them, while the overview said
every module was current ([issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md)).
- A module crash-looping at start — 338 restarts — was found by reading a container's logs by hand.
Nothing on the bus said the mesh's graph had stopped learning.
- A run-once step that fails now stops an upgrade by design ([ADR 0133](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md)),
and the same silence would cover it: the version simply would not appear.
**And the one thing the control plane does emit is refused by its own account.** Answering a catalogue
that asks to catch up, it publishes each recorded build under `mesh.mod.control-plane.event.…` — a
module namespace for a module that does not exist. Its own permissions refuse it, so a catalogue that
restarts gets nothing and keeps its gap. The control plane has facts to state and nowhere to state
them.
## Decision
**The mesh emits the deploy half of the pipeline as facts on the bus.** What a machine now runs, and
what it refused to run and why. Both are facts about the mesh doing its work, in the same form as every
other fact on the bus, so anything that wants them subscribes the way the catalogue subscribes to
`built`.
**The control plane states them, as the holder of the `mesh-controller` seat.** Its facts live under the
seat's own namespace, which is where a role's events belong
([ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md),
[ADR 0129](0129-a-seat-carries-the-protocol-of-its-role.md)) and which survives the control plane being
replaced. That is also what gives the catch-up replay a subject it may publish instead of an invented
module namespace.
**Emitted when what a machine runs changes, not on every convergence pass.** A host reconciles
continuously and reports each time; a fact per pass would be a fact per minute per machine that says
nothing. The report carries the declaration it applied and what changed, so the control plane has what
it needs to speak only when there is something to say.
**A refusal is a fact with a subject in it** — which machine, which resource, and the reason as the host
gave it. A refusal that names only the machine is the silence this record is about, one level up.
**Reports stay where they are.** A node's report remains control traffic that only the control plane
reads. The deploy facts are derived from it, which makes them second-hand on purpose: one emitter, one
ordering, and no widening of the narrowest account in the mesh.
## Options considered
1. **Leave it as it is, and let whatever cares ask the control plane.** Rejected: asking for a fact that
already arrives is the shape the mesh removed everywhere else, and nothing can react at the moment a
machine changes — which is exactly when a graph, an audit or an operator wants to know.
2. **Each node emits its own facts.** Rejected: it widens every host's account to an event namespace, and
a host's authority is deliberately the narrowest in the mesh. Its report already reaches the one thing
that can speak for it.
3. **Widen who may read the control branch.** Rejected: that branch carries what machines say *to* the
control plane, enrolment included. Widening its readers widens that too, for an unrelated reason.
4. **A registry of deploy hooks** — something registers interest and is called. Rejected: an event is
already the mechanism; there is nothing to register, and a callback is an address the mesh spent
[issue 102](../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)
learning not to keep.
5. **Put the facts on the node's own declaration stream.** Rejected: that stream is last-per-subject by
design, so a machine away for an hour gets exactly the current declaration and nothing older. A
history of what happened cannot live in a stream built to forget.
## Consequences
**The audit logger gets the deploy half for nothing**, because it consumes everything.
**A failure becomes visible where the mesh is watched** rather than where someone happened to be
looking. That answers the open question [issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md)
left about a record the store refused.
**The catch-up replay stops being a burst of events.** With the control plane able to state its own
facts, replaying history as if it were happening now is a choice rather than the only option — and the
better shape is the question the catalogue is actually asking, answered once
([design 33](../03-DESIGN/01-to-be/33-the-tools-the-mesh-answers.md)).
**The facts are second-hand.** The control plane says what a machine reported, so a machine that cannot
reach the bus produces no fact at all. Absence is not health, and what a machine was last heard from
stays the place that says so.
**The events stream carries more.** Bounded by emitting on change rather than on every pass, and each
fact is small; the stream's own limits remain what keeps it finite.
## How this is checked
- **The control plane's grant names exactly the subjects it emits**, derived from its seat like every
other principal's, and the composed user list is compared against a golden file — so a fact it cannot
publish fails a test rather than a catalogue's replay.
- **A convergence that changed nothing emits nothing.** A test with two identical reports and one
expected fact, because the failure this guards against is a fact per minute per machine.
- **A refusal names its resource.** A test where a host reports a failed resource and the emitted fact
carries which one and why, not merely that something went wrong.
- **What the mesh emits is what something consumes.** The subject a module declares it consumes derives
to the subject the control plane publishes — the same agreement test that already keeps the
controller's own subscriptions honest.
## References
- [ADR 0041](0041-events-are-a-relationship.md) — an event is a relationship, not a call
- [ADR 0126](0126-a-module-declares-its-own-seats.md) — an event is addressed to its emitter, because the emitter's identity is the meaning
- [ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md), [ADR 0129](0129-a-seat-carries-the-protocol-of-its-role.md) — a role's events belong to the role
- [ADR 0083](0083-one-push-leaves-the-mesh-consistent.md) — a report is held for the store rather than lost
- [ADR 0133](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md) — the gate whose failure this makes visible
- [`03-DESIGN/01-to-be/32-what-a-module-declares.md`](../03-DESIGN/01-to-be/32-what-a-module-declares.md) §6 — the lifecycle, which ends today at a report nobody else may read
- Measured 2026-09-28: 45 minutes of builds recorded nowhere with the overview reporting health; a module at 338 restarts found by hand; the control plane's only emitted event refused by its own permissions
@@ -0,0 +1,153 @@
---
topic: what runs on it
status: accepted
date: 2026-09-28
deciders: jochen
reconstructed: false
supersedes: 0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md
---
# 135. A module version prepares its state before it runs
## Context
[ADR 0133](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md) settled who runs a
module's migrations and when, and it said so in the wrong vocabulary. It put the declaration on a
*container* — "a container may declare steps to run before it" — and derived the scope of the work from
the *machine*. Both are wrong at the level a module author works at, and the second is wrong on the
facts.
**A container is one resource kind the host applies.** A module has code, state and a version; whether
its artifact is an image, a bundle or something later is the mesh's business. The module-facing
vocabulary for a module's own code already exists and has nothing to do with a container runtime: a
module declares **entrypoints** — this file is my tools, this file is my provisioner — and the mesh runs
them. A manifest that says "run this container with these arguments, and here are the volumes and
environment again" has an author writing down the machine's business twice.
**And the scope is not the machine's to decide, because the mesh already decided what a state is.** A
consumer is a module *on a machine* (migration 0015, from
[issue 022](../04-ISSUES/022-one-credential-per-node-per-provision-not-per-module/00-report.md)):
the mesh derives a login per consumer and the provider creates a database owned by exactly that login
([ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)). So a module on three machines is three
consumers, three credentials and three databases. There is no shared state for two machines to race over,
and ADR 0133's central caveat — that a module's migrations must take a lock because two machines might
migrate at once — describes a situation the mesh does not currently produce.
That correction makes the whole "level" question HAL answered with stages disappear: the scope of
preparation is the scope of the state, and the mesh knows it.
What the earlier record got right and this one keeps: the module owns the work, the mesh owns the moment,
the gate is the guarantee, migrations stay forward-only, and none of it can be inferred from inside an
artifact. What produced it also stands — the control plane was replaced with a build carrying a migration,
nothing applied it, and for three quarters of an hour every build was refused by the store with one line
that reached only whoever was waiting on a reply
([issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md)).
## Decision
**A module version declares an entrypoint that prepares its state.** One name in the manifest, in the
same vocabulary as the entrypoints it already declares for its tools and its provisioner. No container,
no command line, no environment, no mounts — those are how a machine runs the module's code, and the
module already said that once.
**The mesh runs it as it runs that module's own code, to completion, in the module's own context.** Every
binding, credential and setting the module's code would receive, because it *is* the module's code. How a
machine does that is the host's business and stays there: for an image artifact it is the step
[ADR 0052](0052-a-step-that-runs-once-before-a-container.md) already defines, and a later kind of artifact
changes the host, not the manifest.
**Preparation gates the version.** A version whose preparation did not succeed does not run — anywhere.
Since the rollout already sends machines one at a time and stops at the first that does not take a
version, a preparation that fails stops the rollout there, leaving every other machine on the version
that works.
**Preparation is scoped to the state, and the mesh derives that scope.** State the mesh provisions is per
consumer — a module on a machine — so preparation happens once per consumer. State the module keeps on
the machine is per machine, which is the same answer. A module that holds an exclusive seat has one of
itself, so its preparation happens once by definition. No level, no election, no cross-node ordering, and
no lock obligation invented for a race the mesh does not create.
**Once per version per state.** A version bump attempts preparation once against each state it has; the
module's own runner decides there is nothing to do, which is what a runner with a version table does
anyway. A retry after a partial failure runs it again, so the work is the module's to make safe against
that — the one obligation no design can remove.
**Forward-only and additive.** Preparation runs while the previous version is still serving, so a
migration that removes or renames what the old code reads breaks the mesh in the window between the two.
**Declared, never inferred.** The control plane cannot see inside an artifact, so a module that ships
migrations and declares no entrypoint is not refusable at registration. It breaks on its first upgrade,
and this record says so rather than implying a check that cannot exist.
## Options considered
1. **A container declares steps before it** — [ADR 0133](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md).
Superseded, not because the mechanism is wrong but because the *declaration* is in the wrong place: it
makes every module author restate the machine's arrangement, and it ties a module's own lifecycle to
one resource kind. The host-side mechanism it named is retained and is now an implementation detail.
2. **Each module prepares itself when it starts** — what the catalogue does today. Rejected: a schema
failure becomes a crash loop rather than a stop, nothing in the declaration says the module has a
state to prepare, and the version serves the moment it starts rather than after the state is right.
3. **The mesh applies migrations itself**, with a driver and a version table per store type. Rejected:
the mesh would have to know one store from another, hold another module's credentials and reach a
machine with them, which [ADR 0005](0005-the-node-host.md) forbids. It is also what forces a stage
system: something central has to decide where the work happens.
4. **A hook lifecycle** — pre-build, post-build, pre-deploy, post-deploy. Rejected: a declaration is a
desired state reconciled forever, so there is no deploy moment to hook. "Pre-deploy" is exactly this
record; pre- and post-build are what a recipe and the artifact list already are; "post-deploy" names
nothing that happens.
5. **A declared level** — once per module, or once per assignment. Rejected: the mesh already knows what a
state is, so asking an author to choose is asking them to restate a fact the mesh holds, with a chance
of contradicting it.
6. **Record a preparation level per module in the store.** Rejected for the reason ADR 0133 gave and this
record keeps: gating makes the invariant true by construction, and a level is a second account of the
same fact.
## Consequences
**An author's whole contract is one line, once.** Write the migration in the module's code, name the
entrypoint that runs it, and every later version rolls out as: build, prepare, run — with nothing
per-version to remember and nothing about the machine to restate. That is the property this exists for.
**Three hand-written steps in the catalogue collapse**, and the control plane's own migrate step stops
repeating its server's environment and mounts.
**The catalogue's self-preparation becomes the exception to remove.** One shape, and the mesh's own
control plane is not an exception either.
**A module scaled across machines with one shared state is not expressible**, and this record does not
make it so. The mesh gives each consumer its own state; a deliberately shared one is a different
provision model, and the place the "once, mesh-wide" question would genuinely return. Named here so it is
a decision when it happens rather than a surprise.
**There is still no readiness-gated step.** Only an action carries `verify`; nothing declares that a
service answers, so preparation that must happen *after* something is serving — seeding through its own
API — remains unexpressible.
**Genesis keeps its own action.** At birth there is no control plane to derive anything, which is what
[ADR 0067](0067-genesis-is-a-pivot.md) says about that moment.
## How this is checked
- **The composition carries the preparation, in the module's own context.** A test on a node's composed
declaration: a version declaring a preparation entrypoint is preceded by it, and what it is given
equals what the module's own code is given — asserted equal rather than written twice, which is the
drift the superseded shape invited.
- **A preparation that fails stops the version.** The host does not go past a step that did not complete,
and the rollout stops at the first machine that did not take a version. Both are existing behaviours
with existing tests; the test for preparation asserts the two together — the machine does not run it,
and the machines after it are left alone.
- **Once per version per state.** A test that a second convergence of the same version prepares nothing,
and that a new version prepares again.
- **The mesh's own control plane declares one.** The case that failed on 2026-09-28 is the case the tests
cover, rather than a case a comment says is covered.
## References
- [ADR 0133](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md) — what this supersedes, and why
- [ADR 0052](0052-a-step-that-runs-once-before-a-container.md) — the host-side step that implements it for an image artifact
- [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md), [issue 022](../04-ISSUES/022-one-credential-per-node-per-provision-not-per-module/00-report.md) — a consumer is a module on a machine, which is what makes the scope derivable
- [ADR 0018](0018-a-picture-is-read-from-what-runs.md) — a digest is the record that something happened
- [ADR 0005](0005-the-node-host.md) — the control plane decides and never touches a machine
- [ADR 0134](0134-the-mesh-says-what-it-applied.md) — what makes a failed preparation visible
- [issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md) — the failure that produced both records
@@ -0,0 +1,106 @@
---
topic: what runs on it
status: accepted
date: 2026-09-28
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md
---
# 136. A step gates its module, not the machine
## Context
[ADR 0052](0052-a-step-that-runs-once-before-a-container.md) made a run-once container a step the host
runs to completion, and gave it the same reach a failed action has: it stops everything the declaration
places after it. When the only steps on the mesh were a broker's seed and a forge's admin account, that
reach was invisible — the thing after the step was the container the step existed for, in the same
module.
[ADR 0135](0135-a-module-version-prepares-its-state-before-it-runs.md) made a step something the mesh
derives for **any** module that prepares its state, and that turns the reach into a fault. A module
whose database is briefly unreachable now stops every module declared after it on that machine, for as
long as it is unreachable.
**The host already rejected this for every other shape, and says why in its own loop.** From
[issue 011](../04-ISSUES/011-one-broken-module-blocks-every-other/00-report.md):
> It used to stop at the first one, and that made one broken resource hold the whole machine hostage: a
> module declaring a package that does not exist meant every module ordered after it was never applied,
> for ever, and the mesh reported "failed" without saying that the rest had not been tried. A machine
> with one bad module and nine good ones ran none of the nine.
Everything is attempted and every failure reported — except an action and a run-once step, kept as the
deliberate exceptions. So the mesh has two rules about the same question and the wider one is now
reachable by any module that declares a schema.
**And it deadlocks a case the catalogue already named.** The catalogue migrates its own schema when it
starts rather than in a step, and says why in its code: *a schema step that had to reach the provider
over the overlay would block the very apply that brings the overlay up*. With a machine-wide gate that
is exactly right — the step fails, the apply stops, the overlay module after it is never applied, and
the next reconcile is blocked the same way. The module that most obviously wants a step could not have
one.
## Decision
**A step gates its own module.** A run-once container that does not complete stops the rest of *that
module's* resources and nothing else. Every other module on the machine is attempted, as every other
shape already is.
**An action still gates the machine.** Genesis is a row of actions, each making the next possible, and
they belong to no module — there is nothing narrower for their reach to be.
**What was not attempted is reported, not inferred from silence.** A skipped resource appears in the
machine's account of the apply as skipped, with the reason, because "not attempted" and "nothing to do"
are different answers and only one of them is somebody's to fix.
**A module is the part of a resource's identity before the first dot**, which is how the mesh composes
them. What the mesh declares in its own right — a guard, an opening, the adoption's own resources —
belongs to no module, and its gate is therefore the machine's.
## Options considered
1. **Leave the reach as it is.** Rejected: it reintroduces, through a mechanism now derived for every
module, exactly the fault issue 011 removed. A mesh where one module's unreachable database stops a
machine converging is worse than one where that module alone is behind.
2. **Make preparation not a gate at all** — run it and carry on. Rejected: then a version serves against
a state nobody shaped, which is the whole of what ADR 0135 exists to prevent.
3. **Order every module's step before everything else on the machine**, so a gate stops nothing that
matters. Rejected: it inverts the order a module needs — its files and directories are declared before
its step because the step reads them — and it would still stop later modules.
4. **Let a module declare how far its step reaches.** Rejected: the answer is the same for every module,
and a field would let one be wrong about it.
## Consequences
**The catalogue can move to a step.** The reason it migrates at start — that a step blocks the apply
that would make its provider reachable — stops being true: the step fails, that module waits, the
overlay comes up, and the next reconcile prepares it. One shape for the whole mesh, which is what
ADR 0135 asked for and could not have had.
**A module can sit behind while the machine is otherwise current.** That is the honest state and it is
what the report now says. It also means a preparation that never succeeds is a module that never
upgrades, quietly, until somebody reads the report — which is an argument for
[ADR 0134](0134-the-mesh-says-what-it-applied.md) rather than against this.
**A module's resources must be ordered within the module for the gate to mean anything.** They already
are: the mesh composes a module's resources in the order its manifest declares them, and its own
workload comes after the files it reads.
## How this is checked
- **A failed step stops its module and nothing else.** A test with two modules: the one whose step
failed does not start its workload, the other starts, and the error still says the failure gated
something. It fails against the previous behaviour, which is how it was written.
- **An action still stops the machine.** The existing test for a failed action is unchanged, and a step
with no module in its identity — which is what genesis carries — takes the same path.
- **The report names what was skipped.** Asserted in the same test, because a gate nobody can see is
indistinguishable from a module that had nothing to do.
## References
- [ADR 0052](0052-a-step-that-runs-once-before-a-container.md) — the step this narrows
- [ADR 0135](0135-a-module-version-prepares-its-state-before-it-runs.md) — what made the reach reachable
- [issue 011](../04-ISSUES/011-one-broken-module-blocks-every-other/00-report.md) — the same fault, removed once already
- [ADR 0134](0134-the-mesh-says-what-it-applied.md) — how a module left behind becomes visible
- mesh-host `internal/apply` — the loop whose own comment argued this case for every other shape
@@ -0,0 +1,117 @@
---
topic: what runs on it
status: superseded
date: 2026-09-28
deciders: jochen
extends: 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
reconstructed: false
superseded-by: 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md
---
# 137. A machine says which networks it routes
## Context
The filter the mesh derives denies forwarding by default, because without a forward chain it says
nothing about a container's published port
([ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md), [issue 047](../04-ISSUES/047-the-firewall-does-not-cover-published-container-ports/00-report.md)).
To keep a machine's own containers working it then allows two ranges: the container runtime's
default bridge pool, and the pool its compose files are given. Those two are named in the
controller's code, with a comment saying what the gap is:
> A machine whose runtime is configured with something else needs this to say so — which is a thing
> the mesh cannot derive and a reason this list is named here rather than computed.
**There was no way to say so.** The list was a constant. A machine whose guests live anywhere else
was filtered by a rule that looked deliberate and was a guess.
**Measured, on the day a workstation was converged.** Flipping it cut egress for five of its
container networks at once, and for every network its test beds create — the beds allocate a fresh
range per run, from a pool neither default covers. Nothing reported a fault. The containers could
not reach anything, the machine went on reporting that it had applied what it was told, and the
converge preview had said nothing about it either, because the preview lists what *listens* and
routing is not a listener.
**And two questions, not one.** A guest also asks its host for an address and for names. Both arrive
at the input chain, where nothing declared them, so denying by default left the guests of a routed
network with no address and no resolution — which is not a closed port but a network that does not
function, asked for by this machine's own guest.
**Why the machine cannot simply be read.** A test bed creates its bridge while it runs, between one
declaration and the next, so a filter derived from what the machine last reported would be correct
only for the networks that already existed when it was composed. A declared range covers the ones
that do not exist yet.
## Decision
**A machine says which networks it routes for what it hosts, and the filter forwards them.** A
node-level fact, beside the node's public domain
([ADR 0066](0066-public-routing-is-name-agnostic.md)) and for the same reason: the
machine routes them, and the module that loads the filter holds a seat and may be replaced.
**Added to the runtime's defaults, never replacing them.** A machine that names one range has not
stopped hosting whatever was already on the runtime's own pools, and replacing would trade one
silent breakage for another.
**Their guests keep address and name service.** For a network that was named, the input chain admits
that network's own DHCP and DNS, and nothing else: everything else a guest might want from its host
is a port somebody declares, like every other port on this machine.
**Said in CIDR form and checked when it is said.** An entry that does not parse is a line nftables
refuses, and a refused ruleset is a machine filtering nothing while its unit reports a fault — so
the refusal happens where a person can read it, not on the machine.
**A machine that says nothing is filtered exactly as before.** Every machine already converged is
untouched by this.
## Options considered
1. **Leave it constant and edit the code per installation.** Rejected: the value is a property of
one machine, the code is the whole mesh's, and the two ranges as they stand describe a machine
whose runtime was left at its defaults. It is also how this got here.
2. **Derive it from what the machine reports.** Rejected as insufficient, not as wrong: it cannot
cover a network created between two declarations, which is precisely the case that was broken. It
would also make the filter follow whatever appeared on the machine, which is a firewall that
widens itself.
3. **A per-node setting on the module that loads the filter.** Rejected: the machine routes the
networks. The filter module holds a node-scoped seat and is meant to be replaceable, and a
replacement must not lose the machine's own truth.
4. **Replace the defaults with what is said.** Rejected: see the decision. The first machine to name
its bed range would lose its containers.
5. **Admit all input from a routed network, not only address and name service.** Rejected: that is
every port on the machine open to anything it hosts, which is the derivation abandoned.
## Consequences
**The converge preview says what a machine routes**, including when it routes nothing but the
defaults, with the command that changes it. The preview's own sentence about traffic it cannot
preview stays, because a tunnel and the found firewall's NAT are still not previewable.
**A machine whose guests are already broken by an earlier flip is fixed by saying its networks and
pushing**, with no flip to undo.
**The list is one more thing that can be wrong and stale.** A range removed from the machine and
left here keeps forwarding for a network that no longer exists, which admits nothing, because there
is no guest on it to admit. That is the safe direction of being out of date.
## How this is checked
- **What a machine says it routes is forwarded, and its guests keep address and name service.** A
test renders a ruleset for a machine that names one range and asserts both chains, per chain body
so a line in the wrong chain cannot pass it. It fails against the previous behaviour, which is how
it was written.
- **The runtime's own defaults survive naming a range.** Asserted in the same test.
- **A machine that names nothing renders byte-identically to one that names nil**, so every machine
already behind this filter is untouched.
- **Each family is matched in its own syntax.** A test with one v4 and one v6 network asserts
`ip saddr` and `ip6 saddr`, because one set holding both is a syntax error and a ruleset that does
not load is a machine filtering nothing.
- **An entry that is not a network is refused where it is said**, by the parse in the setter.
## References
- [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md) — the derived filter this completes
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — the precedent for a node-level fact
- [issue 047](../04-ISSUES/047-the-firewall-does-not-cover-published-container-ports/00-report.md) — why there is a forward chain at all
- [issue 137](../04-ISSUES/137-converging-a-machine-cut-off-its-own-guests/00-report.md) — the measurement that produced this
- mesh-controller `internal/catalogue/filtering.go` — the constant whose own comment named this gap
@@ -0,0 +1,179 @@
---
topic: what runs on it
status: accepted
date: 2026-09-28
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md
---
# 138. An assignment binds an endpoint and says how far it reaches
## Context
[ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) settled that a machine's
packet filter is derived from what its modules declare they listen on, and that the `from` of a
listen "is the whole of public-versus-internal". That was true of the packet filter, and it turned
out to be true of nothing else.
Reachability is now settled three times, in three places, by three mechanisms that cannot disagree
out loud ([issue 140](../04-ISSUES/140-an-endpoints-reach-is-not-declared/00-report.md)):
- **The filter** reads a listen's source, and a per-node setting may override it. That setting has
exactly one caller in the control plane — the function that builds the node's rules.
- **The names** come from a route contribution, which names a label and a port and says nothing
about reach. The reverse proxy composes a **public** name and an **internal** name for every route
it is given, because it can.
- **The certificate authority** follows from which names exist. Measured on the control-node: an
identity provider carries a public certificate valid 90 days and an internal one valid 24 hours and
renewed daily. No assignment asked for either.
So *this endpoint must not be public* cannot be written. It is therefore enforced by nothing, while a
public certificate for that very name is obtained automatically — the fault
[how-we-build.md](../00-META/how-we-build.md) names, an unenforced rule being indistinguishable from
a wrong one, with the additional cost that the wrong thing is done eagerly.
And a port that is not routed cannot be spoken about at all beyond the filter. The forge serves git
over ssh; that endpoint has no name, no certificate and no way to be called public except a key only
the filter reads.
**Two per-node settings already exist and are half of this.** One gives a module's declared port a
machine port. One overrides a declared port's source. They key on port numbers, so nothing ties a
port to the route that serves it: a route contribution names a port too, and the two are equal only
by coincidence.
**Where this belongs is already decided.** [ADR 0046](0046-a-module-configuration-is-its-assignments-not-its-manifest.md)
says a module's configuration is its assignments. Whether the forge answers git-over-ssh from the
public internet is a fact about one installation and one machine, not a property of the software —
and [ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md) already refuses an
installation's decisions in a definition.
## Considered Options
1. **Leave reach in the manifest, as `from` today.** Rejected: it is an installation's decision
written into the definition, and it cannot differ between two machines running the same module —
which is exactly the case the forge presents.
2. **Extend the existing source override to the names and the certificate, without naming
endpoints.** Rejected: it keys on a port number. A module's route contribution names a port as
well, and nothing says the two are the same thing, so one statement cannot be made to reach all
three mechanisms. Naming the endpoint is what makes that possible.
3. **Derive reach from whether the node has a public domain recorded.** Rejected: that is a property
of the machine, and two endpoints on one machine differ — a database and a web front end on the
same host.
4. **A fourth reach for "public name, internal authority"** — a name that resolves publicly and must
not appear in a public issuance log, obtained by DNS-01. Deferred, not rejected: it is a real case
and it is a question about which challenge an authority uses, not about how far an endpoint
reaches. Left to the certificate work as an open question.
5. **Make the manifest silent on reach and require every assignment to state it.** Rejected for the
transition: every endpoint reachable today would close until an assignment named it, which is a
flag day across the whole catalogue.
## Decision
**A module declares named endpoints.** An endpoint is one port the module serves, with a name the
module chooses, its protocol, and what it is for. A route contribution **names the endpoint it
routes** rather than repeating a port number. The manifest says what the module serves and what it
would serve it to by default; it does not say what this installation does with it.
**An assignment binds each endpoint and says how far it reaches.** Per node: the machine port the
endpoint is published on, and its **reach** — one of `internal`, `public` or `both`. An assignment
that states nothing keeps the manifest's default, so no machine changes until an assignment says so.
**Reach means all three mechanisms at once, and is the only thing that decides them.**
- `internal` — the filter opens the machine port to the private network; the proxy serves the
internal name and not the public one; the certificate comes from the mesh's own authority.
- `public` — the filter opens it to anywhere; the proxy serves the public name; the certificate
comes from the public authority.
- `both` — both names, each from its own authority, and the filter opens to anywhere.
**An endpoint that is not routed is reached but never named.** An endpoint with no route contribution
yields filter rules and nothing else: no name is composed and no certificate is requested. Git over
ssh is that case, and it is the case the model could not express.
**The authority stops being chosen by which names happen to exist.** The proxy composes the names the
assignments asked for, and asks each name's own authority for it. A name nobody asked for is not
composed, so it is not certified.
**The two existing settings are this, completed.** The per-node port mapping becomes the endpoint's
binding. The per-node source override becomes its reach, widened from the filter alone to the names
and the certificate as well.
## Progressive insight — 2026-09-29, from building it
**Reach does not mean the same thing to the filter for an endpoint the proxy serves.** The decision
above says `internal` means "the filter opens the machine port to the private network" and `public`
means "the filter opens it to anywhere". For a routed endpoint the second half is wrong, and
[ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) already said so before this
record was written: *a public service is exposed through the proxy, not by opening its own port* — it
listens `from: mesh`, only the proxy reaches it, and it is exposed by name.
Found by trying to express one real module, not by review. Its routed name must be public, because
browsers post to it; its machine-side port must not be, because that port serves the dashboard in
cleartext. Under one value driving both, saying "public" would have reopened a port an operator had
just closed. Measured the same evening: that module's routed name answered from the internet over TLS
while its machine-side port was refused from the same place. The port is not the path.
So the reach of a **routed** endpoint asks for names, and its port keeps what the manifest said. The
reach of an **unrouted** endpoint — git over ssh, a mail port, the bus — governs the port, because
there is no name and the port is the only way in. That is the same split this record already draws in
*an endpoint that is not routed is reached but never named*; what it got wrong was carrying the filter
across it.
This corrects a fact, not the decision: one statement per endpoint, three things derived from it and
none of them deciding on its own, all stand. The table in the decision should be read with the filter
column applying to an unrouted endpoint.
## Consequences
- **A manifest gains endpoint names, and a route contribution names an endpoint instead of a port.**
Every routed module's manifest changes. The word ships one release before any manifest uses it, and
reaches the build machine and the control plane first.
- **One derived value is read by three things** — the filter's rules, the proxy's contributions, the
certificate request — so they can no longer disagree, and a disagreement becomes a refusal at the
assignment rather than a surprise on a machine.
- **A name that must not be public becomes writable, and therefore checkable.** It also gives
[issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md) a
declared answer to read: which endpoints are internal is what says whose root must be installed
where.
- **[Issue 139](../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md)
becomes answerable**: the endpoint's assignment names the machine that serves it, which is the fact
the internal name should be composed from.
- **Reach becomes reportable.** The mesh can say, per endpoint, where it is reachable from and which
authority holds its certificate — neither of which `status` can say today.
- **This narrows [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md).** Its
decision stands: the firewall is derived and host-applied, not a provider. What no longer holds is
that a listen's `from` is the whole of public-versus-internal; it is the filter's share of a
statement that also governs names and certificates.
- **What got harder:** every endpoint needs a name, including a module that serves exactly one port
and had no reason to name it. And an installation that wants a module public must now say so on the
assignment rather than inheriting it from the definition, which is more to say and the reason it is
right.
## How it is checked
- **One module, two endpoints, different reach.** A module declaring an internal endpoint and a
public one renders a filter opening one to the private network and one to anywhere, asserted per
chain body so a rule in the wrong chain cannot pass.
- **The names follow the reach.** The same module's routed endpoint composes the internal name only
when internal, the public name only when public, and both when both — and a certificate is
requested from the matching authority for each name composed and for no other. This fails against
the previous behaviour, where both names and both certificates are always composed, which is how
it is written.
- **An unrouted endpoint is filtered and never named.** Asserted for an endpoint with reach and no
route contribution: rules rendered, no contribution, no certificate request.
- **An assignment naming an endpoint the module does not declare is refused where it is said**, as is
a reach that is not one of the three — before it reaches a machine, because a ruleset that does not
load is a machine filtering nothing.
- **An assignment that states nothing renders byte-identically to today**, so every machine already
converged is untouched until its assignment says otherwise.
## References
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — narrowed here
- [ADR 0046](0046-a-module-configuration-is-its-assignments-not-its-manifest.md) — where reach belongs
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — the public name this composes
- [ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md) — why reach is not a definition's
- [issue 140](../04-ISSUES/140-an-endpoints-reach-is-not-declared/00-report.md) — the measurement
- [issue 139](../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md),
[issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md)
@@ -0,0 +1,137 @@
---
topic: what runs on it
status: superseded
date: 2026-09-28
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0137-a-machine-says-which-networks-it-routes.md
superseded-by: 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md
---
# 139. A network is forwarded because a module declared it
## Context
[ADR 0137](0137-a-machine-says-which-networks-it-routes.md), decided the same week, gave a machine a
way to say which networks it routes for its guests. It was written because the derived filter's
forward chain allowed two ranges named as constants in the control plane's source — the container
runtime's bridge pool, and part of the pool its compose files are given — with a comment admitting
the gap: *a machine whose runtime is configured with something else needs this to say so, which is a
thing the mesh cannot derive.*
**It can be derived, and from the right place.** Measured on the last machine still to be converged
([issue 141](../04-ISSUES/141-the-forward-chain-does-not-follow-the-modules/00-report.md)): twenty-one
container networks, nine inside the runtime's bridge pool, twelve in the other private range, and six
of those outside the constant's lower bound — so the flip would have cut their guests off exactly as
it did on the workstation that produced 0137.
Naming a range to cover the six is what 0137 provides for, and it is the wrong instrument. Of those
six networks, **four are networks the mesh's own modules declare**, present as network resources in
the node's plan and created by the host because a module asked for them. **Two are the predecessor's
leftovers** — compose networks of services the mesh does not run. Any range wide enough to keep the
four forwards the two as well: a firewall widened by hand to protect networks that should not exist.
The mesh already knows which of the twenty-one are its own, because it made them.
**And the node's configuration is meant to follow the modules assigned to it.** That is the mesh's
founding shape — the machine runs modules, and its files, its filter and its accounts are composed
from what runs there ([ADR 0005](0005-the-node-host.md),
[ADR 0010](0010-delivery.md),
[ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)). The forward chain is the
one derived thing that consults a constant and a list a person types.
## Considered Options
1. **Keep 0137 as it stands** — two constants plus a named list. Rejected: the list is written in
addresses, and addresses are what the runtime allocates, so the only entry safe enough to keep a
machine working is wider than the truth. It cannot distinguish a network the mesh made from one
left behind, which is the distinction that decides whether forwarding it is correct.
2. **Derive it from what the machine reports.** Still rejected, on 0137's own grounds: a test bed
creates its bridge between one declaration and the next, and a filter that follows whatever
appeared on a machine is a firewall that widens itself. **This decision is not that** — see below.
3. **Have the control plane allocate each module network's range from a pool it owns,** so it can
render the address itself. Rejected: more machinery for no gain. The runtime already allocates and
the host already knows, and taking allocation over means the mesh owning an address space it has no
other reason to own.
4. **Have each module declare its network's range.** Rejected by
[ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md): a definition names no address,
and the same definition runs on machines whose runtimes have allocated differently.
## Decision
**A network is forwarded because a module declared it.** Per node, the forward chain forwards the
networks of the modules assigned there, and by default nothing else. A module unassigned stops being
forwarded at the next reconcile.
**The host resolves a declared network to its addresses.** A network resource carries a name; the
runtime allocates the subnet when the network is created. So the control plane declares *forward the
networks these modules asked for* and the host — which made them, and already resolves a container by
its name — renders the addresses. [ADR 0005](0005-the-node-host.md) holds: the host applies, it does
not decide.
**Deriving from the declaration is not deriving from the machine.** Both of 0137's objections fall
away. The set is known before the network exists, because a module declared it, so a network created
between two declarations is already in the one that asked for it. And it cannot widen itself: a
network nobody declared is never forwarded, however it appeared on the machine.
**The runtime's own default bridge is forwarded, from what the runtime reports.** Containers that name
no module network attach to it, and it belongs to the runtime rather than to any module — so the host
renders it from what the runtime says, not from a range named in the control plane. The constants go.
**What a machine says is for guests no module declares.** A test bed is not a module and its range is
not a module's; that is the case 0137's mechanism is for, and it keeps it — added to the derived set,
never replacing it, as 0137 decided. Narrowed to that, it is named for it.
**Their guests keep address and name service**, per declared network, unchanged from
[ADR 0137](0137-a-machine-says-which-networks-it-routes.md): the input chain admits that network's own
DHCP and DNS and nothing else.
## Consequences
- **The two constants are removed**, and with them the class of fault that a machine's guests depend
on a range that describes some other machine.
- **This is a behaviour change, not a refactor.** On the machine measured, the derived set and the
constant do not cover the same ground — that is the whole reason for the record. A machine whose
module networks happen to fall inside the old ranges renders the same rules.
- **A range that exists only to keep a leftover alive becomes visible as such**, because it will not
be in the derived set and has to be said out loud to survive.
- **`node networks` narrows** to guests no module declares, and the preview says which of a machine's
networks are the mesh's and which are not, so the difference is readable before a flip rather than
after.
- **A module's declaration gains nothing.** It already declares its network; what changes is that the
filter reads it.
- **What got harder:** the host renders part of the forward chain from what it created, so the
control plane no longer holds the whole rule set as text. The rule the mesh states is the set of
networks; the addresses are the machine's.
## How it is checked
- **Only declared networks are forwarded.** A node with two modules that declare networks renders
forward rules for exactly those two, and none for a third network present on the machine that no
module declared. This fails against the previous behaviour, which forwards by range and cannot tell
them apart, and that is how it is written.
- **Unassigning a module removes its network's rule** at the next reconcile, asserted on the rendered
chain rather than on the intent.
- **Guests of a declared network keep address and name service**, asserted per chain body so a line in
the wrong chain cannot pass — carried from 0137.
- **The runtime's own default bridge comes from the runtime**, asserted by rendering for a runtime
whose default bridge is somewhere other than the range the constant named.
- **A machine that names a range for guests no module declares still gets it**, added to the derived
set and not replacing it.
- **Each family is matched in its own syntax**, carried from 0137: one set holding both is a syntax
error, and a ruleset that does not load is a machine filtering nothing while its unit reports a
fault.
## References
- [ADR 0137](0137-a-machine-says-which-networks-it-routes.md) — narrowed here; its mechanism keeps the
case it is right for
- [ADR 0005](0005-the-node-host.md) — the host applies; the addresses are the machine's
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the filter is derived from
what runs there
- [ADR 0112](0112-a-module-definition-names-no-node-mesh-or-path.md) — why a module does not name its
range
- [issue 141](../04-ISSUES/141-the-forward-chain-does-not-follow-the-modules/00-report.md) — the
measurement
- [issue 137](../04-ISSUES/137-converging-a-machine-cut-off-its-own-guests/00-report.md) — the
breakage that produced 0137
@@ -0,0 +1,146 @@
---
topic: what runs on it
status: accepted
date: 2026-09-28
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md
supersedes:
- 02-DECISIONS/0137-a-machine-says-which-networks-it-routes.md
- 02-DECISIONS/0139-a-network-is-forwarded-because-a-module-declared-it.md
---
# 140. The filter constrains what arrives from outside, and says nothing about a machine's own guests
## Context
The filter the mesh derives blocks traffic passing *through* a machine unless something allows it,
because a container's published port is traffic passing through rather than traffic arriving at the
machine itself ([ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md),
[issue 047](../04-ISSUES/047-the-firewall-does-not-cover-published-container-ports/00-report.md)).
Having blocked all of it, the filter then had to let the machine's own containers reach outward again.
It does that by listing the address ranges those containers sit on.
As rendered on a converged workstation today:
```
policy drop
ct state established,related accept
ip saddr 172.16.0.0/12 accept
ip saddr 192.168.128.0/17 accept
ip saddr 10.0.0.0/8 accept
ip saddr 192.168.16.0/20 accept
... four more
```
Two of those ranges were constants in the control plane's source. The rest were typed by the operator
after [ADR 0137](0137-a-machine-says-which-networks-it-routes.md), which existed to make the typing
possible, because converging that workstation had cut every one of its containers off from the
internet and nothing reported a fault
([issue 137](../04-ISSUES/137-converging-a-machine-cut-off-its-own-guests/00-report.md)).
**The list is the mistake, not its contents.** Every attempt to make it correct fails the same way.
A constant describes one machine. A typed range goes stale, and cannot tell a network the mesh made
from one a predecessor left behind — measured on the control-node, where six such ranges fall outside
the constants and two of the six belong to services the mesh does not run
([issue 141](../04-ISSUES/141-the-forward-chain-does-not-follow-the-modules/00-report.md)).
[ADR 0139](0139-a-network-is-forwarded-because-a-module-declared-it.md) tried to generate the same
list from the modules and put half the rule set on the machine to do it. Three records, one list, and
the list should not exist.
**Because the mesh has no policy about a container reaching outward.** What the filter is for is
stated in [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md): which port is open,
and to whom. That is about what arrives. A container of this machine's own opening a connection to
something else is not a port being opened to anybody, and enumerating the addresses it might do so
from is bookkeeping about the machine's internal plumbing, which the mesh neither owns nor can know.
**The system being replaced never had this fault, and its rule says why.** The chain still protecting
the control-node applies only to traffic arriving on that machine's outward link, and leaves
everything else alone. The mesh's filter dropped that distinction and replaced it with a list of
addresses.
## Considered Options
1. **Keep the list and generate it better** — from the modules' declared networks, or from what the
machine reports. Rejected: [ADR 0139](0139-a-network-is-forwarded-because-a-module-declared-it.md)
is that, and it puts part of the rule set on the machine, which makes the rule set partly the
machine's and the derivation advisory.
2. **Name the guest links instead of their addresses, and allow only those.** Rejected as more than is
needed: it fails in the safe direction, but it is still a list that has to keep up with the
machine, and the thing it protects against — a container reaching outward — is not a thing the mesh
has a position on.
3. **Do not block traffic passing through at all.** Rejected: that is
[issue 047](../04-ISSUES/047-the-firewall-does-not-cover-published-container-ports/00-report.md),
where a published port was reachable from anywhere because no rule mentioned it.
4. **Constrain what arrives from outside, and nothing else.** Adopted.
## Decision
**The filter constrains traffic arriving from outside the machine, and says nothing about traffic that
did not.** Traffic passing through the machine is allowed unless it arrived on one of the machine's
outward links, in which case it is allowed only where a declared endpoint's reach admits it
([ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md)). A container of this
machine's own reaching anywhere is not filtered, because the mesh has no position on it.
**A machine says which of its links face outside.** One node-level fact, reported by the machine the
way it already reports the kind of firewall it found and the tunnel it carries — not a setting, not a
list of addresses, and not something anybody types. It does not change when a module is added or
removed, which is what separates it from the list it replaces.
**A machine that has reported no outward link is sent no filter.** Rendering a rule around a link
whose name is not known produces a rule set that does not load, which is a machine filtering nothing
while its unit reports success. The refusal happens in the control plane, where a person reads it, and
the machine keeps the filter it already has.
**No addresses of the machine's own networks appear in the filter.** The two constants are removed and
`node networks` is removed with them, along with everything any machine was told to say through it.
Ports continue to follow the modules exactly as before: a module assigned to a machine opens the port
its assignment says it reaches on, and nothing about a network is said anywhere.
## Consequences
- **Three records collapse into one rule.** 0137 and 0139 are superseded. What 0137 was right about —
that converging a machine had silently cut off its own containers, and that nothing previewed it — is
answered by removing the cause rather than by giving the operator a way to compensate for it.
- **Every machine already converged loses its declared ranges and keeps working**, because the traffic
those ranges allowed is now allowed by not having arrived from outside. The workstation's five ranges
and the laptop's one are deleted rather than migrated.
- **A machine's test beds stop being a special case.** A bed's network is created while the machine
runs and was the case no list could cover; it is now covered by not being mentioned.
- **A new fact travels in the report**, and the control plane refuses to compose a filter without it,
so the order of the roll-out matters: the machines report before the control plane depends on it.
- **A machine with more than one outward link says so**, and a machine that acquires one while the mesh
is not looking is treated as internal until its next report. That window is the cost of this shape;
it is bounded by the report interval, and it exists on machines whose outward link changes, which
are the machines with nothing published to the outside.
- **What got harder:** nothing in the declaration, and one more thing a machine must be able to work
out about itself. A machine that cannot say which link faces outside cannot be given a filter.
## How it is checked
- **A machine's own container reaches outward with no network named anywhere.** A bed converges a
machine carrying containers on several networks, none of them mentioned in any setting, and each
reaches out afterwards. This fails against the previous behaviour, where the same flip cut them off,
and that is how it is written.
- **A port declared reachable from outside is reachable; one that is not, is not.** Probed from off the
machine's private network, for a published port and for an undeclared one, before and after the flip.
- **A network created after the filter was composed needs no new filter.** A network is made on the
machine after its last declaration and a container on it reaches out, with nothing re-sent.
- **No address of a machine's own networks appears in a rendered filter**, asserted on the text so a
range cannot creep back in.
- **A machine that reports no outward link is sent no filter, and the refusal names it** — asserted in
the control plane, and that the machine's existing filter is left alone.
- **A machine reporting two outward links has both constrained**, asserted per chain body so a rule
covering one and not the other cannot pass.
## References
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — what the filter is for
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — what admits traffic
arriving from outside
- [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md) — why traffic passing through is
filtered at all
- [ADR 0137](0137-a-machine-says-which-networks-it-routes.md),
[ADR 0139](0139-a-network-is-forwarded-because-a-module-declared-it.md) — superseded here
- [issue 137](../04-ISSUES/137-converging-a-machine-cut-off-its-own-guests/00-report.md),
[issue 141](../04-ISSUES/141-the-forward-chain-does-not-follow-the-modules/00-report.md)
@@ -0,0 +1,171 @@
---
topic: what runs on it
status: accepted
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0005-the-node-host.md
---
# 141. The host delivers its own successor, and versions live side by side
## Context
[Issue 142](../04-ISSUES/142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md). A
merge builds every changed module and the control plane — which is itself a module — and the result
reaches the machines running it with nobody asking. The host is the exception: it is not a build
target, no declaration delivers it, and every machine in this mesh runs a byte-identical binary that
somebody built on a workstation and copied out.
The half that *recovers* from a bad host exists. `internal/upgrade` can tell that the executable this
process started from was replaced on disk, and it records which version last completed a reconcile.
The launcher counts consecutive failed starts, calls a rollback at the limit, and treats a clean exit
as the host standing aside so that the next loop runs whatever is on disk now. That supervision is
complete and correct.
Two things make it dead code:
- **`Replaced()` is called by nothing but its own tests.** Nothing tells the running host that a
successor is waiting.
- **The rollback resolves a version through the machine's package manager** — `pacman -U` from the
package cache. No machine here has the host installed as a package, so the recovery cannot run on
any of them; and being written in one package manager's terms, it cannot run on two of the three
operating systems the host is built for — [ADR 0005](0005-the-node-host.md) builds one binary per
operating system, pinned at link time.
**The record already points at the answer.** What is kept is a *version*, not a path. Keeping a
version is only useful to something that can choose between versions present on the machine, which is
what the package manager was being asked to do. The versions can simply be on disk.
## Considered Options
1. **Deliver the host as a package, as the rollback assumes.** Rejected: it needs a package built and
a repository trusted per operating system, three of each, and the existing `package` resource
asserts presence and deliberately never a version — "version is the package manager's business and
the mesh does not hold a second opinion about it" — so it cannot ask for a particular host anyway.
Heaviest of the three and the only one that is different on every machine.
2. **Write the new binary over the running one.** Rejected on a fact: a running executable cannot be
truncated, and `archive` opens what it unpacks with `O_TRUNC`. It could be made to write and
rename, which is better hygiene and worth doing for its own sake, but it buys nothing here that
option 3 does not, and it leaves rollback with nowhere to go back to.
3. **Versions side by side; the newest retires the old.** Adopted.
## Decision
**A host version is delivered as an archive into a directory named for it, and never over a running
one.** The declaration names it like any other archive — fetched by digest, the digest checked before
anything is unpacked. Nothing new travels, no new resource kind, and no change to how archives are
applied, because the path being written is not the path being executed.
**The launcher starts the most recently delivered version.** That is what "the newest" means: the
version whose directory arrived last. It reads no pointer and follows no link — the mesh creates no
links ([ADR 0012](0012-the-mesh-creates-no-symlinks.md)) — and the version is in the path, so nothing
has to be told what is running.
**The running host stands aside for a successor, and only between reconciles.** Finding a newer
version delivered, it finishes the reconcile it is in and exits cleanly. The launcher already reads a
clean exit as exactly this and starts what is on disk now. A host that stood aside mid-apply is the
half-configured machine this project exists to prevent, so the check happens at the boundary and
nowhere else.
**A version that completes a reconcile records itself, and retires what came before it.** The
known-good record is written as it is today. Then versions older than the one before the running one
are removed: the running version and its predecessor are kept, which is exactly what a rollback
needs, and nothing else accumulates.
**Rollback starts the previous version instead of reinstalling a package.** At the failure limit the
launcher pins the known-good version and starts that, once. The second failure is still a different
diagnosis — the previously working version does not run either, so it is the machine and not the
binary — and the halt is unchanged. No package manager, no package cache, and the same script on every
operating system.
**A machine says which host version it is running,** on the report it already sends, beside the other
facts it states about itself. Without it nothing can say a machine is behind, so "every machine
current with its source" cannot include the host.
## Progressive insight — 2026-09-29, the same day
**The delivery is not "nothing new", and this record said it was.** The decision above stands and is
built: versions side by side, the newest runs, the running host stands aside between reconciles, a
completed reconcile retires what is older than the predecessor, rollback picks a directory. What was
wrong was a claim about how a version reaches a machine. The paragraph on delivery said the
declaration "names it like any other archive… nothing new travels, no new resource kind"; the second
half is true and the first is not, because two things the delivery needs do not exist:
- **Nothing can compile it.** A `bundle` artifact is compiled by a closed list of toolchains —
typescript and python — whose own comment says adding a language is a decision, because a language
used by *modules* needs an SDK carrying the broker client, the event envelope and tool serving. The
host uses none of that: it is what applies modules, not one of them. So the obligation that list
warns about attaches to a module written in a language, not to the language being buildable, and
the control plane — also written in Go — is built as an image from a Dockerfile rather than through
a toolchain at all.
- **A version cannot reach the path.** An `archive` resource names a fixed path in the manifest, and
nothing interpolates the built version into it, so nothing can ask for
`…/versions/<version>/`.
Neither changes what was decided, which options were weighed, or any consequence: the shape is
unaffected and the host half is merged and tested. What it changes is the cost, which this record
understated as none. The remaining work is a way to build the host and a way to name a version in a
path, and until both exist nothing delivers a version and every machine takes the fallback — which is
what every machine does today.
> **Progressive insight — 2026-09-30. Both of those exist now.** The paragraph above named two missing
> things and they are built: a Go toolchain, based on a new `mesh-tools-go` module so the compiler is
> named and not pinned, and `${version}` in any value of a resource that uses an archive or a bundle.
> The mesh compiles its own host and publishes it to its own registry, measured — a statically linked
> stripped binary, fetched back out and run. **The cost was larger again than this note said**: three
> more things in the path assumed one language or one shape, and a fourth was in the base image.
> The account is [issue 142](../04-ISSUES/142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md).
>
> The version in a path is the artifact's **digest**, not the commit this note's own wording would
> suggest. Two builds of one commit are meant to be the same bytes, so a content-addressed version
> means an unchanged build keeps the path it had; a commit-named one would move for an identical binary
> and recreate everything reading it.
>
> **Still nothing delivers a version to a machine.** The host is a module and builds, and declares no
> resources, so the bundle sits in the registry and no machine is asked to take it. That is the next
> piece, and the decision above is unchanged by any of this.
## Consequences
- **The host becomes a build target and a module** — a module whose resource is the next host, applied
by the host that is running. The bootstrap is not circular because the two are different versions in
different directories.
- **Rollback becomes usable on every machine**, having been usable on none. It also stops being
written in one operating system's terms.
- **One copy by hand remains, once.** The first host that understands versioned directories cannot be
fetched by a host that does not. That copy is the last, and it is the honest cost of the change
rather than a step in the design.
- **Two versions occupy disk instead of one.** About nine megabytes. The predecessor is the price of a
rollback that does not depend on a cache somebody else may clean.
- **What got harder:** a host must now be able to find its own successor and to judge when it is safe
to stand aside. Both are between reconciles, which is the only moment the host is not mid-change.
- **A machine that is never told a newer version keeps running what it has**, indefinitely and
visibly, because its report says which version that is.
## How it is checked
- **A delivered version is run, and the old one is not.** A bed delivers a second version to a machine
running the first; the host exits between reconciles, the launcher starts the new one, and the
machine reports the new version. This fails against the previous behaviour, where nothing notices a
delivered version at all.
- **It stands aside between reconciles and never inside one.** Asserted by delivering a version while
an apply is in flight: the apply completes, and the exit follows it.
- **A version that will not start is rolled back to its predecessor, once**, and the second failure
halts with the machine named rather than the binary — asserted with no package manager involved.
- **A completed reconcile retires what is older than the predecessor**, and never the predecessor
itself, because that is what a rollback needs. Asserted on the directory afterwards.
- **The report names the running version**, asserted end to end rather than on the function that reads
it, since the point is that the control plane can tell a machine is behind.
- **The launcher picks the newest delivered version** with no pointer file and no link, asserted by
delivering two and checking which runs.
## References
- [ADR 0005](0005-the-node-host.md) — the host, and what its supervision is for
- [ADR 0010](0010-delivery.md) — a declaration is owned resources; this adds no kind to it
- [ADR 0012](0012-the-mesh-creates-no-symlinks.md) — why the version is in the path
- [ADR 0005](0005-the-node-host.md), *it is built per operating system* — why a rollback written in
one package manager's terms was wrong for two of three
- [issue 142](../04-ISSUES/142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md) — the
measurement
@@ -0,0 +1,158 @@
---
topic: the mesh
status: accepted
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0141-the-host-delivers-its-own-successor.md
---
# 142. The mesh delivers its own components as binaries, not as container images
## Context
Measured on the control-node, 2026-09-29:
| what | how it runs | publishes |
|---|---|---|
| host | a binary on the machine | — |
| controller, catalogue, builder, vault | containers | nothing |
| store, registry, broker | containers | ports |
**The mesh's own software is delivered two ways, and the difference is not a property of the
software.** The host and the controller are both written in the same language, both the mesh's own,
both doing the mesh's own work. One is an image fetched from a registry. The other is a file somebody
copied to four machines, owned by no package, built by nothing
([issue 142](../04-ISSUES/142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md)).
**The reason is not a judgement about either, it is that images are the only delivery that works.**
There is no way to put a binary on a machine. The host is hand-copied because of that, and the
controller is an image because of that. Neither was chosen on its merits.
What it costs, all of it measured rather than argued:
- **Genesis must raise a container runtime before the control plane can exist.** The bundle carries
three images and one of them is the controller, *"in the bundle for the same reason they are: there
is nothing to fetch it with yet"*
([design 07](../03-DESIGN/01-to-be/07-the-foundation.md)). So the hardest moment in the mesh's life
has a prerequisite that the thing being started does not need.
- **Updating the control plane depends on the control plane.** Its image is fetched from the registry,
which is a container the controller manages.
- **A change to the host cannot be rolled out at all.** Every machine here runs a byte-identical
hand-copied binary. A change merged yesterday reached none of them.
- **Compiling the language the mesh is written in is not a capability of the builder.** The bundle
toolchains are typescript — real, with a registered base module — and python, which is named in the
list and absent from the catalogue. The controller is built as an image from a Dockerfile, which is
the per-repository incantation the bundle toolchain exists to abolish
([design 18](../03-DESIGN/01-to-be/18-building-a-module.md)).
The half that *receives* a binary safely is already built and tested
([ADR 0141](0141-the-host-delivers-its-own-successor.md)): versions side by side in directories named
for them, the newest run, the running one standing aside between reconciles, retirement keeping the
predecessor, and a rollback that chooses a directory. What is missing is everything that puts one
there.
## Considered Options
1. **Leave it as it is.** Rejected: it is not a design, it is the reach of one mechanism. And it is
what makes a host change undeliverable.
2. **Containerise the host too**, so everything is delivered one way. Rejected: the host is what
starts the container runtime and what applies containers. A host in a container is the bootstrap
problem made total, and the machine would have no way back from a bad one.
3. **Deliver the mesh's components as operating-system packages.** Rejected for the reason
[ADR 0141](0141-the-host-delivers-its-own-successor.md) rejected it for the host: a package and a
trusted repository per operating system, three of each, and the `package` resource asserts presence
and deliberately never a version.
4. **Binaries for the mesh's own components, containers for third-party software.** Adopted.
## Decision
**The mesh's own components are delivered as binaries on the machine.** The host, the controller, the
catalogue, the builder, the vault — the software this project writes. They are delivered by the
mechanism [ADR 0141](0141-the-host-delivers-its-own-successor.md) built: an archive, fetched by
digest, unpacked into a directory named for its version, with the running one standing aside between
reconciles and a rollback that chooses the predecessor.
**Third-party software stays a container.** The store, the registry, the broker. They are somebody
else's build, they are already adopted as modules
([ADR 0078](0078-the-store-and-broker-are-modules.md)), and an image is the right way to carry
somebody else's software. **The container runtime remains required** — modules use it — so this
removes a dependency from the control plane, not from the machine.
**The builder compiles the languages the mesh is written in.** A toolchain for Go, with a base module
providing the compiler, exactly as typescript has. The obligation the toolchain list warns about — an
SDK carrying the broker client, the envelope and tool serving — attaches to a *module* written in a
language, not to the language being compilable. None of these components is a module in that sense;
the host is what applies modules.
**An artifact says what it targets.** A compiled binary is per operating system, pinned at link time
([ADR 0005](0005-the-node-host.md)), and a toolchain deliberately takes nothing from the module,
because anything a module could override there it would be writing a Dockerfile to override. So the
target is a property of the artifact rather than of the recipe, and one artifact declared per target
is one build each.
**A component's version comes from where it sits, not from its linker.** It is unpacked into a
directory named for its version, so it can read its own version from its path. The stamp goes, and
with it the need for a build to know what it will be called.
**Genesis carries a binary reference where it carried an image reference.** The principle does not
change — the bundle names a thing by digest and the host fetches it, pinned because nothing can
resolve a version when no mesh exists — and the container runtime stops being a prerequisite for the
control plane. It stays a prerequisite for the store and the broker, which is where it belongs.
**The order is staged, and each step stands alone.** Compiling Go; an artifact naming its target;
delivering a binary; the host as the first component delivered; the controller, catalogue, builder and
vault out of their containers; genesis last. Genesis is last for the reason it is always last: it
matters for a machine nobody has yet, and every earlier step is provable on a mesh that exists.
## Consequences
- **One delivery for the mesh's own software**, so a change to the host ships the way a change to the
controller does, and neither is copied by hand.
- **The control plane stops depending on a container runtime and on its own registry.** Both remain on
the machine for other reasons; neither gates the control plane's own life any more.
- **`Replaced()`, the known-good record and the launcher's rollback stop being dead code.** They were
written for this and have been called by nothing but their tests.
- **Four more components gain a rollback they do not have.** Today a bad controller image is recovered
by an operator; under this it is recovered the way a bad host is.
- **Two versions of each component occupy disk.** Around nine megabytes each. The predecessor is what a
rollback needs.
- **Genesis gets smaller, not larger.** One fewer image to carry and one fewer runtime to raise before
the control plane.
- **This does not make the components smaller or simpler.** They are the same programs; what changes is
how they arrive. A reader expecting the containers to have been hiding complexity will not find any.
- **What got harder:** the builder gains a language, artifacts gain a target, and the mesh gains a
second kind of thing it must deliver correctly — one where getting it wrong takes the control plane
down rather than a module. That is why the host is first: it is the component whose recovery is
already built and tested.
## How it is checked
- **A component is delivered and runs, with nothing copied by hand.** A bed builds the host from its
repository, delivers it to a machine running an older one, and the machine reports the new version.
This fails today at the first step, because nothing builds it.
- **Each target is built once and only the matching one is delivered.** Asserted by declaring an
artifact per operating system and checking that a machine is offered the one it can run — a host
built for another is what ADR 0005's link-time pin exists to refuse.
- **A component reads its version from its path**, asserted by unpacking the same bytes into two
differently named directories and seeing each report its own.
- **A bad component is rolled back without an operator**, for the host first: a version that will not
start is replaced by its predecessor once, and the second failure halts naming the machine.
- **The control plane comes up with no registry reachable**, which is the dependency this removes —
asserted by raising it with the registry stopped.
- **Genesis raises a control plane with no container runtime running**, and raises the store and the
broker afterwards. Last, and on a machine with nothing on it.
- **A published port count that does not change.** The mesh's own components publish nothing today, so
moving them out of containers must not open anything — asserted on the machine's reachable set before
and after, which the converge preview already reads.
## References
- [ADR 0141](0141-the-host-delivers-its-own-successor.md) — the receiving half, already built
- [ADR 0005](0005-the-node-host.md) — the host, its supervision, and one binary per operating system
- [ADR 0078](0078-the-store-and-broker-are-modules.md) — why third-party software stays a container
- [ADR 0006](0006-the-substrate-and-the-control-plane.md) — what genesis must raise, and in what order
- [issue 142](../04-ISSUES/142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md) — the
measurement that started this
- [design 07](../03-DESIGN/01-to-be/07-the-foundation.md) — the bundle's three images, one of them the
controller
@@ -0,0 +1,135 @@
---
topic: what runs on it
status: superseded
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0010-delivery.md
superseded-by: 02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md
---
# 143. A consumer verifies the grant it is given
## Context
A **grant** is what the mesh writes on a consumer's machine so it can reach a provider. The real one
the forge receives for its database, as it arrives:
```
provision postgres-database
at <the provider's machine, by name>
port the machine port the provider is published on
as the role the provider created for this consumer
```
with the credential sealed in a separate file. Four facts and a password, and they are the whole
mechanism by which anything in the mesh reaches anything else.
**The mesh asserts that claim and never finds out whether it is true.**
[Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md):
converging a machine dropped the path from a container to a port on its own machine, and for eleven
hours the mesh answered *all doing what they were told, all heard from, every module current with its
source* while a web application logged, six thousand times:
```
connection to server at "<the machine>" (10.10.0.1), port 6852 failed: timeout expired
```
Every check the mesh makes passed, because every check it makes is about the relationship between the
mesh and a machine: the declaration was applied, the digest matched, every container named was running.
None of them asks whether a consumer can reach what it requires — though the mesh composed the grant
and therefore knows the consumer, the machine, the address, the port and the credential.
**And where the check runs decides whether it catches anything.** The rule in force admitted the
machines' own addresses on the private network. A dial from the *machine* to its own address carries
exactly such a source address, so a check run by the host on its own behalf would have matched that rule
and passed — while every container on the machine was refused. This is inference from the rule that was
loaded, not a measurement: the fault was found and fixed before anyone thought to dial from the host.
It is enough to decide the question, because a check whose position differs from the consumer's is
testing something nobody asked about.
## Considered Options
1. **The control plane dials each provision.** Rejected, and it is the tempting one because the control
plane holds every fact. It sits on the provider's machine for most provisions here and reaches the
address by a path no consumer uses; in the measured outage it would have passed throughout.
2. **The host dials on the consumer's behalf, from the machine.** Rejected for the reason above: the
machine's network position is not the consumer's, and the one outage this exists to catch is exactly
a difference between them.
3. **Ask the module.** Rejected: a module is arbitrary software that the mesh does not write. Some could
report on their provisions and most cannot, and a check that covers the modules that opted in tells
nobody anything about the rest.
4. **Read the module's logs.** Rejected: the failure was in a log the whole time, and reading a module's
logs makes the mesh depend on the wording of software it does not control.
5. **The consumer verifies it, from its own network position.** Adopted.
## Decision
**A consumer verifies each grant it is given, from its own network position.** After a reconcile has
applied a grant, the machine opens a connection to the address and port that grant names, from inside
the consumer's own network namespace — the same position the consumer's software dials from, which is
the only position that answers the question the grant asks.
**It is a connection, not a conversation.** Whether the port accepts a connection is what a grant
claims; whether the credential is right, the role exists or the schema is current is the provider's to
answer and the consumer's to discover. A check that spoke each provision's protocol would be a second
implementation of every provision, and would fail for reasons that are not the mesh's.
**One failure is not news.** A provider restarting is ordinary, and so is a consumer between containers.
A grant is reported unreachable only after it has failed on **consecutive** reconciles, and the count is
what the machine reports rather than the last attempt — so a reader can tell "it was briefly away" from
"it has never worked".
**A grant that cannot be checked is said to be unchecked, never assumed good.** A consumer that is not
running has no network position to dial from; that is not a broken grant and must not read as one. It is
also not a verified grant, and the two are different sentences.
**What it costs to be wrong is the constraint on all of it.** A check that reports a working provision
broken trains a reader to ignore the report, which is worse than having none — the fault this
repository keeps finding, one level up. So the threshold is consecutive failures, the check is the
cheapest thing that answers the question, and an unknown is reported as unknown.
**The mesh says it where it says everything else.** A machine's report carries its unreachable grants,
and `status` names them beside what is out of date — so "every module current with its source" stops
being the whole of what the mesh will tell you about a machine whose modules cannot reach each other.
## Consequences
- **The mesh can be wrong out loud.** It has been able to assert a grant and not check it; now a grant
that does not work is a thing the mesh says, and the eleven hours of issue 145 become minutes.
- **The host gains the ability to act from a container's network position**, which it has not needed
before. That is a real capability and the only one this needs.
- **A machine reports something that is not about the declaration.** Everything it reports today is
what it applied and what it holds; this is the first thing it says about whether what it applied
works.
- **A provision with no port is not checked**, because there is nothing to dial. Several are files and
secrets, and saying "checked" about those would be the appearance of verification that this record
exists to remove.
- **What got harder:** a reconcile does more than apply. Every grant adds a connection attempt on a
cadence, which is cheap individually and worth naming: a machine with many consumers dials once per
grant per reconcile.
## How it is checked
- **The outage is caught.** A bed drops the path from a consumer's network position to a provider's
port while leaving the machine's own path to it open — the exact shape of issue 145 — and the grant
reads unreachable. This fails against the previous behaviour, where nothing reported anything, and
against a check run from the machine, which passes while the consumer cannot reach it.
- **A restarting provider is not an outage.** One failed reconcile reports nothing; the count rises and
falls, and the grant reads reachable again without anybody acting.
- **A consumer that is not running reads unchecked, not broken**, asserted separately from the
unreachable case because they are different sentences.
- **A provision with no port is not claimed to be checked.**
- **The report carries the count, not the last attempt**, so "briefly away" and "never worked" are
distinguishable by a reader who sees only the report.
- **`status` names an unreachable grant**, asserted on the output, since a check nothing surfaces is
the same as no check.
## References
- [ADR 0010](0010-delivery.md) — the declaration is owned resources; a grant is one of them
- [ADR 0009](0009-modules-and-the-graph.md) — what a provision and a consumer are
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
— the eleven hours
- [issue 136](../04-ISSUES/136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md) — the
same distance between a declaration and a machine, one level down
@@ -0,0 +1,121 @@
---
topic: what runs on it
status: accepted
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md
supersedes: 02-DECISIONS/0143-a-consumer-verifies-the-grant-it-is-given.md
---
# 144. Anything on a machine may call anything on it, and that is the whole of "local"
## Context
Everything in the mesh should be able to call:
- what runs on the same machine;
- another machine's service over the private network, if that service is exposed there;
- another machine's service over the public network, if it is exposed there.
Three cases. The filter had two of them.
**The first was broken and the break was invisible.** A service exposed to the private network rendered
as the machines' own addresses on it. A caller on the machine carries such an address; a caller inside
one of that machine's containers carries a bridge address and matched nothing. Measured:
```
the machine: local 10.10.0.1 dev lo src 10.10.0.1
a container: 10.10.0.1 via 172.17.0.1 dev eth0 src 172.17.0.8
```
Same destination, same machine, two source addresses. The rule named the first and silently refused the
second, so a module reaching its database on its own machine's name timed out for eleven hours
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)).
**The second case works, and by accident.** A caller on another machine reaches the private network over
the tunnel, and arrives carrying that machine's own address — so the rule matches. It would not have
matched the caller's own address either; the tunnel rewrites it. That two of three cases worked is why
this looked correct.
**[ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) answered the wrong question.** Written
hours earlier, it proposed that a consumer verify each grant it is given by opening a connection from
its own network position — and it went to some length about *which* position, because whether a caller
sat in a container changed the answer. That difference was the bug. A verification mechanism would have
reported this outage sooner and would not have prevented it, and the machinery it needed existed only
because the rule was wrong. The remedy for a configuration error is the correct configuration.
**And a module is not a container.** A module is software that delivers one or more services, and it may
do that as a container, an installed package with a unit, a binary, or files something else reads. Of 72
modules in the catalogue, 61 happen to use a container and 11 do not — among them the resolver, the ssh
daemon and the intrusion-prevention module. A rule that reasons about containers describes most of the
mesh and not the mesh.
## Considered Options
1. **A line per service admitting the machine's own callers.** Rejected: it is what was written first,
and it only ever covers the services somebody remembered to think about. It also states, service by
service, a thing that is true of the machine.
2. **Verify each grant from the consumer's position** ([ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md)).
Rejected as a remedy: it observes the fault rather than removing it, and the question it agonised over
— which network position — exists only while the fault does.
3. **Enumerate the addresses a machine's callers may have.** Rejected for the reason no address is named
anywhere in this filter any more ([ADR 0140](0140-the-filter-constrains-what-arrives-from-outside.md)):
a range describes one machine and goes stale in silence.
4. **Local is not filtered, stated once.** Adopted.
## Decision
**Anything on a machine may call anything on that machine, and the filter says so once.** Not per
service, not per port, and not by naming who the callers are: traffic that did not arrive from outside
the machine and did not arrive over the private network is the machine's own, and is admitted. It is
asked by the link the traffic arrived on, because that is a fact about the machine rather than a list
that describes one.
**Local is not a boundary this mesh draws.** Whether a caller is a container, a unit, or the operator's
shell changes nothing, because the thing being decided is "is this the same machine" and the answer does
not depend on the form the caller takes.
**The other two cases are unchanged and are now legible beside it.** A service exposed to the private
network admits the machines on it; a service exposed publicly admits anything. Three cases, three lines,
and a reader can see all three at once.
**[ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) is superseded and nothing replaces it.**
Whether the mesh should check that a grant works is a real question — it reported this machine healthy
for eleven hours — but it is a question about what the mesh can say, not about what it should do, and it
must stand on its own rather than as the remedy for a rule that was wrong. It is not built.
## Consequences
- **The three things everything should be able to call are three lines**, and the first is one line
rather than one per service, so a service added tomorrow is reachable locally without anybody
remembering to say so.
- **A form of module stops mattering to the filter.** The 11 modules that are not containers were never
affected by this bug and were never the reason it was hard to see; they are the reason the rule should
never have mentioned containers.
- **The mesh still cannot say when a grant stops working.** That is the live gap, recorded in issue 145
and no longer pretending to have an answer.
- **What got harder:** nothing. This removes a line per service and replaces it with one.
## How it is checked
- **A caller on the machine reaches a service on it, in the input chain**, asserted on that chain's own
body — because the forward chain carries the same line in the same words, and an assertion on the
whole rendered file passed with the input chain's copy deleted. That is what
[ADR 0137](0137-a-machine-says-which-networks-it-routes.md)'s tests already say to do.
- **It is one rule, not one per service.** Asserted by rendering two services of different reach and
refusing a per-port local line.
- **The three reaches render as three lines**, asserted together, so the whole of what the filter says
about who may call what is one test.
- **The measured case:** from a container on the machine, a service exposed to the private network on
that machine answers. This is the outage, and it fails against the rule this replaces.
## References
- [ADR 0045](0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md) — the filter is the sum
of what its modules listen on
- [ADR 0140](0140-the-filter-constrains-what-arrives-from-outside.md) — why no address is named
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — internal and public,
the other two cases
- [ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) — superseded here
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
@@ -0,0 +1,120 @@
---
topic: what runs on it
status: superseded
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md
superseded-by: 02-DECISIONS/0146-connectivity-is-checked-by-name-per-hosting-form.md
---
# 145. A module checks what the mesh claims is reachable, and it checks itself
## Context
The mesh asserts three things are callable ([ADR 0144](0144-anything-on-a-machine-may-call-anything-on-it.md)):
what runs on the same machine, another machine's service exposed to the private network, and another
machine's service exposed publicly. It has never checked any of them.
[Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md):
the first of the three was broken for eleven hours and the mesh answered *all heard from, every module
current with its source* throughout. Every check it makes is about the relationship between the mesh and
a machine — applied, current, containers running — and none about whether anything can reach anything.
**A first answer was drafted and withdrawn.** [ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md)
put the check inside the host, verifying each grant from the consumer's network position. It was
superseded because the difference it worked so hard to reproduce — whether a caller sat in a container —
was the bug itself. What survives from it is the part that was right: a check run from the wrong place
proves nothing, and the mesh's own reports are not evidence about the network.
**The mesh already has the shape for this and it is a module.** A module can declare a container that
runs on a cadence ([ADR 0053](0053-a-step-that-runs-on-a-schedule.md), and three modules already use
`*/5 * * * *`), can be given the mesh's roster as a rendered fact — every machine's name, address and
this node's own identity, the same mechanism the resolver and the operator's ssh configuration use — and
can emit what it found on the bus. Nothing new is needed to build this except the module.
**What it must not check is the trap.** The obvious probe target is ssh: present on every machine, never
closed by design. Dialling it would have passed throughout the outage, because ssh is admitted
unconditionally and the thing that broke was a service exposed to the private network. A checker whose
probe is unconditionally open measures the one path that cannot fail, which is the failure this whole
sequence keeps producing — a check that reads as verification and verifies nothing.
## Considered Options
1. **The host verifies each grant** ([ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md)).
Superseded. It needed the host to act from another network position, which is machinery that exists
only while local calls are filtered wrongly.
2. **The control plane dials every node.** Rejected: it sits on one machine and reaches the others by a
path no ordinary caller uses. It would have passed throughout the outage.
3. **Probe an existing service.** Rejected for the target problem above: the services guaranteed on every
machine are the ones that are never closed, so they cannot fail the way the mesh fails.
4. **A module on every machine that serves its own probe and dials the others'.** Adopted.
## Decision
**A module runs on every machine, serves an endpoint of its own, and dials every other machine's.** The
probe is the module's own endpoint, declared reachable over the private network — so the thing being
dialled is admitted by exactly the rule that governs every other internally-exposed service, and fails
when that rule is wrong. A second endpoint, declared public, does the same for the public path where a
machine has one.
**It checks the three cases the mesh claims, by name:**
- its **own machine**, by dialling its own machine's address — the case that broke, and the only one that
distinguishes a caller on the machine from a caller in one of its containers;
- **each other machine over the private network**;
- **each machine's public path**, where one is recorded.
**It resolves before it dials, and says which failed.** A name that does not resolve and a port that does
not answer are different faults with different owners, and a checker that reports one sentence for both
sends a reader to the wrong place.
**It runs where the callers run.** The module's own code in its own container, on the cadence the mesh
already has, from the same position as every other module on that machine. It is not the host and not the
control plane, and that is the whole point.
**It says what it found and nothing else.** It emits results; it repairs nothing, opens nothing and holds
no credential beyond its own. A checker that fixes things is a second control plane.
**One failure is not a fault.** A machine rebooting is ordinary. A path is reported broken after it has
failed on consecutive runs, and the count travels with the result so a reader can tell "briefly away"
from "never worked" — the one thing [ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) got
right and worth keeping.
## Consequences
- **The mesh gains the ability to be wrong out loud about the network.** Eleven hours becomes two runs.
- **It is a module, so it is assigned, built, pushed and reported on like everything else** — no new host
capability, no new vocabulary, nothing in the control plane that has to know about checking.
- **Its own endpoint is the instrument.** That is what makes it able to fail; it also means the checker
must be assigned to a machine before that machine can be checked, and a machine without it is
unchecked rather than healthy.
- **It cannot check what it cannot be told.** The roster gives it machines; it does not give it every
module's endpoints, so this checks the paths the mesh claims and not every grant in the mesh. That is
the honest scope of a first one, and the difference is worth saying rather than growing quietly.
- **What got harder:** one more module on every machine, and a module whose whole purpose is to fail
visibly when something else is wrong. Its own failures will be read as the mesh's, which is the cost of
an instrument.
## How it is checked
- **It catches the measured outage.** A bed closes the path from a container to a service exposed to the
private network on its own machine — issue 145's shape — and the checker reports its own machine
unreachable while every other path still reads reachable. This fails against a probe on a port that is
never closed, which is the wrong target this record exists to name.
- **A machine rebooting is not a fault**: one failed run reports nothing, the count rises and falls.
- **A name that does not resolve is reported as that**, not as a port that did not answer.
- **It reports and does not act**: asserted by giving it a broken path and checking nothing on the machine
changed.
- **A machine without the module reads unchecked**, never healthy — asserted on what the mesh says about
a machine it is not assigned to.
## References
- [ADR 0144](0144-anything-on-a-machine-may-call-anything-on-it.md) — the three things that must be callable
- [ADR 0143](0143-a-consumer-verifies-the-grant-it-is-given.md) — superseded; what survives is that the
position matters
- [ADR 0053](0053-a-step-that-runs-on-a-schedule.md) — the cadence
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — internal and public,
which the probe endpoints declare
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
@@ -0,0 +1,125 @@
---
topic: what runs on it
status: accepted
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0145-a-module-checks-what-the-mesh-claims-is-reachable.md
supersedes: 02-DECISIONS/0145-a-module-checks-what-the-mesh-claims-is-reachable.md
---
# 146. Connectivity is checked by name, per hosting form, with a valid certificate
## Context
[ADR 0145](0145-a-module-checks-what-the-mesh-claims-is-reachable.md) decided that a module checks what
the mesh claims is reachable, from where the callers are, because the mesh reported four machines healthy
for eleven hours while a module could not reach its database
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)).
That decision stands. What it got wrong is everything about *what* is dialled.
It dialled a raw port on each machine's address. Three things are wrong with that:
- **A raw port is not how anything in this mesh is reached.** A real caller resolves a name, the proxy
answers it, and the proxy reaches the service. A check that dials a port tests the last hop of a path
with four hops in it, and the three it skips — resolution, the proxy, the certificate — are where most
of the mesh's connectivity actually lives.
- **It tested one hosting form.** A module is software that delivers services, and it may deliver them
from a container, from a unit the mesh writes for its own code, or from a unit a package ships. Those
are three different paths to the same machine, and the outage that produced this was two of them
disagreeing. A probe served one way measures one way.
- **It said nothing about certificates.** An internal name that resolves, routes and answers over TLS
that nothing can verify is not a working path; it is a working path for whoever holds the proxy's
trust and nobody else.
## Decision
**Each hosting form gets its own endpoint, its own route and therefore its own name.** On every machine:
| name | what serves it |
|---|---|
| `connect-docker.<node>.internal` | a container |
| `connect-process.<node>.internal` | the mesh's own code, in a unit the mesh writes |
| `connect-unit.<node>.internal` | a unit a package ships |
and the same set under each machine's public domain where it has one — `connect-docker.<domain>` and its
siblings. The names are the instrument: a failure reads as *`connect-docker.g14.internal` did not answer*,
which says which machine and which hosting form without anybody interpreting anything.
**Every machine checks every machine, by name, over TLS, verifying the certificate.** Not a port, not an
address: resolve the name, connect, complete the handshake, check the certificate against the authority
that should have issued it — the mesh's own for an internal name, a public one for a public name. That is
the whole path a real caller takes, and each step failing is reported as itself.
**No name is written anywhere.** The machines come from the roster the mesh already renders as a fact, and
the labels are the module's. A machine that joins appears in every other machine's roster on the next
push, and they begin checking it without an edit.
**And the module arrives on a machine because the machine exists, not because somebody assigned it.** A
machine that joins and does not have it is worse than unchecked: every other machine is already dialling
its names, so it reads as broken everywhere until someone notices. This is the part the mesh cannot
currently express — see below — and it is the part that makes the rest safe.
**What survives from 0145**, unchanged: it reports and repairs nothing; one failure is not a fault and a
path is broken after consecutive runs with the count travelling with the result; findings are said on the
bus, because a finding in a file on the machine is what this exists to end; and the bus is the one path
that cannot report its own failure, so an emit that does not land is written locally and nowhere else.
## What this needs that the mesh does not have
Named here rather than assumed, because each is a decision of its own and this record is not the place to
make them:
1. **A module that every machine has.** `ScopeNode` means *at most one holder per node* — an exclusivity
rule, not an obligation — and nothing assigns a module at enrolment. Today the resolver, the packet
filter, ssh and intrusion prevention are each assigned per machine by hand, which is the same gap
wearing different clothes.
2. **A container running a module's own bundle.** A `process` runs the mesh's own compiled code with no
image; a `container` needs an image of the module's own, which means a Dockerfile — the thing the
`bundle` artifact exists to abolish. Nothing in the catalogue runs a bundle in a container, so
`connect-docker` has no shape yet.
3. **A unit a package ships, for `connect-unit`.** The `service` resource puts an existing unit into a
state and deliberately installs none, so this form needs a package that serves a port — and naming a
program the machine may not have is
[issue 136](../04-ISSUES/136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md).
4. **A machine's public domain in the roster fact.** The fact carries each machine's name, mesh name,
address and operator account. The public names cannot be composed without the domain.
5. **Something that installs the mesh's own root on a machine.** This is
[issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md), open
since before any of this. Until it is closed, every internal name will fail certificate verification
from every machine — correctly, because nothing can verify it. That is the checker working, and it is
worth saying in advance so the first run is not read as the checker being broken.
## Consequences
- **A failure names the machine and the hosting form.** That is the whole gain over a port: eleven hours
became two runs under 0145, and under this it also becomes one line that says where to look.
- **The checker surfaces issue 129 immediately**, and will report every internal name unverifiable until
it is fixed. A reader must be told that before the first run rather than after.
- **Five things must be built before this is what it says it is**, and until they are, what exists is a
port dial from one position — useful, and not this.
- **What got harder:** a module with three hosting forms of the same trivial service is a strange thing to
read. It is justified only because those three forms are how the mesh actually runs software, and a
checker that tested one of them would keep the class of outage it exists to catch.
## How it is checked
- **A name per hosting form answers from every machine**, asserted by name and not by port.
- **A certificate that does not verify is reported as that**, distinctly from a name that does not resolve
and a port that does not answer — three faults, three owners.
- **A machine that joins is checked by every other machine without an edit**, asserted by adding one to a
bed and looking at what the others dial on their next run.
- **A machine that joins has the module**, which is gap 1 above and is the assertion that cannot be
written yet.
- **The measured outage is still caught**: the path from a container to a service on its own machine is
closed and `connect-docker.<that node>.internal` fails from that machine while the others still pass.
## References
- [ADR 0145](0145-a-module-checks-what-the-mesh-claims-is-reachable.md) — superseded; its core stands
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — the two reaches these
names come from
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — a label plus a domain, which is why no name is written
- [issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md) — what the
internal names will fail on until it is closed
- [issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
@@ -0,0 +1,169 @@
---
topic: what runs on it
status: accepted
date: 2026-09-29
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md
---
# 147. A module anchors the mesh's authority on a machine, and takes it away again
## Context
The mesh runs its own certificate authority and every internal name is served with a certificate
from it. No machine trusts it. On an enrolled, adopted workstation — on the private network,
resolving through the mesh's resolver — every internal HTTPS name fails verification with
*unable to get local issuer certificate*
([issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md)).
The certificates are genuine; nothing on the machine has ever been told what issued them.
The authority's only consumer today is a proxy, which fetches the root into a directory of its own
and hands it to one program ([ADR 0098](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md)).
That is enough for the proxy and for nothing else: a browser, `git` over HTTPS, `curl`, a package
manager and every module that calls another module by an internal name read the machine's trust
store, which holds the predecessor's authority and a developer tool's local root, and nothing of
the mesh's.
The predecessor wrote its root into every machine it set up. Removing it was deliberate — an
honest failure beats a name that verifies for the wrong reason — and it leaves the mesh with no
answer at all until this one lands. It is also what keeps the predecessor alive on the machines
that still speak TLS to a mesh name.
**What makes this a decision rather than a patch** is where the knowledge goes. Two mechanisms in
the mesh already write things onto a machine because it is on the private network: `/etc/hosts`
and the registry's plaintext trust ([ADR 0082](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)).
Following that precedent, the controller would inject an anchor into every such machine's
declaration, and issue 129 proposed exactly that. It would work. It would also put *where this
operating system keeps trust anchors* and *which command refreshes its bundles* into the control
plane, for a fact the control plane does not have (the root does not exist until the authority has
run) and a machine that may have no reason to verify a mesh name at all.
## Considered Options
1. **The controller injects the anchor into every machine on the private network**, the
`/etc/hosts` and insecure-registry shape. Rejected: being on the network is what makes the
registry reachable, and that is why network presence is the right trigger *there* — the trust
and the reachability are the same fact. Trusting an authority is not the same fact as being
able to reach it, and the anchor's path and the bundle refresh are a property of the machine's
operating system, which is the host's half of the mesh, not the controller's.
2. **A new host primitive — a `trust-anchor` resource type.** Rejected for now, not on principle.
The host's vocabulary should grow when a shape cannot be said with what exists, and this one
can: a file and a service already express it, as the packet filter proves
([ADR 0140](0140-the-filter-constrains-what-arrives-from-outside.md), whose module writes a
unit file and a service and nothing else). The primitive becomes right the moment a second
operating system is in play, because the anchor directory and the refresh command are exactly
the difference `internal/system` exists to hold. Until then it would be a vocabulary word with
one speaker.
3. **A module that requires the authority, fetches its root, installs it as an anchor and
refreshes the machine's bundles — and removes both when it is no longer assigned.** Adopted.
## Decision
**A machine trusts the mesh's authority because a module put its root there, and stops trusting it
when that module is taken away.**
1. **The module requires `internal-acme-ca`** and reads the provider's bound address and the path
it serves its root at. It requires nothing else and provides nothing: it is a consumer of the
authority like any other.
2. **It fetches the root over the mesh's own network, without prior trust**, because there is no
prior trust to have — this is the module that establishes it — and the network is what
authenticates the fetch ([ADR 0098](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md),
the same reasoning that lets the proxy fetch it). What it accepts is checked: a body that is
not a certificate fails, and the failure is the module's, not a later handshake's.
3. **It installs the root where this machine's TLS clients look, and refreshes the extracted
bundles** — the command that does the refresh is an ordinary part of the unit that places the
anchor, not a new thing the mesh can be asked to do.
4. **Removal is symmetric and is the same unit's business.** Undeclared, the host stops the unit;
stopping it removes the anchor and refreshes the bundles again. A machine that leaves the mesh
stops trusting the mesh, without anybody remembering to go and look.
5. **It is an ordinary assignment.** No machine is given it automatically. A machine that verifies
a mesh name is assigned it, and a machine that does not is not — which is the same statement
the mesh already makes about every other module, and is why this is not the controller's
business.
**One operating system, said out loud.** The anchor directory and the refresh command in the
module today are Arch's. On a machine that is not Arch the unit fails, visibly, rather than
writing a file nothing reads. That is the accurate failure, and it is the signal that option 2
above has become right.
## How this is checked
- **The verification that could not succeed before.** On a machine holding the module, a plain
client fetches an internal HTTPS name with no `-k` and no bundle argument and verifies. On a
machine without it, the same fetch fails with *unable to get local issuer certificate*. Both
halves, because only the pair distinguishes "the anchor works" from "something else already
trusted it".
- **The removal half, in the same bed:** unassign the module, refetch, and the failure returns.
Checking only the arrival is how a trust store fills up with authorities nobody can account for.
- **What is deliberately not checked here:** that the authority issues, that a name resolves, that
the proxy serves. Those have their own beds, and this module's bed passing for those reasons is
the failure mode this record is most exposed to — which is why the negative half is not optional.
**What this bed is dialled at, and why it is the authority itself.** The authority serves its own
API with a certificate it issued, so the handshake under test needs nothing else in the mesh to be
right. A trust bed that reached for a routed name through the proxy would be passing or failing for
the proxy's reasons and the resolver's.
**Written, and not yet run** *(2026-09-29)*. The bed is `trust-anchor` in the lab, and it cannot
execute: raising a foundation fails before any module is reached, in both bundles that exist
([issue 146](../04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md)).
So what stands behind this record today is the rendering — the script the machine would run names
the authority it was bound to, checked in the control plane's own test suite — and **not** a machine
that verified anything. That is a weaker thing than the paragraph above describes, and it stays
written this way until the bed runs.
> **Progressive insight — 2026-09-30. It has now been run, on the live mesh rather than in the bed.**
> The paragraph above said nothing had verified anything, and something has. The module was registered
> from the catalogue, assigned to a workstation, and checked in the form this section prescribes — the
> authority's own API, so the handshake needs nothing else in the mesh to be right:
>
> ```
> $ curl -sS -o /dev/null -w '%{http_code}' https://<the authority>:9000/health
> 200
> subject=CN=Step Online CA
> issuer=O=Mesh Internal CA, CN=Mesh Internal CA Intermediate CA
> Verify return code: 0 (ok)
> ```
>
> **Both halves.** Unassigning and pushing removed the anchor, emptied the trust store of the mesh's
> authority, and returned the plain client to *unable to get local issuer certificate* — then assigning
> again restored it. The negative half is what distinguishes the anchor working from something else
> having trusted it, and it is the half nothing had ever exercised.
>
> **One thing this found that is not in the module.** The removal only works because the *host* removes
> the service before the script: stopping the unit is what deletes the certificate and refreshes the
> bundles, and it needs the script it calls to still exist. Nothing in the module states that ordering;
> the symmetry this record claims rests on it.
>
> Run on the live mesh because that is where a change is verified now
> ([ADR 0149](0149-the-live-mesh-is-the-test-bed.md)), and the bed still cannot raise a foundation. The
> evidence is [issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/02-resolution.md).
> Extended the same day to every converged machine — `novox`, `g14` and `shanks` each hold the anchor
> and verify with a plain client. `ace` is excluded on purpose: it is adopted, so a module assigned
> there is held rather than run, which is right and is not trust.
## Consequences
The predecessor's authority can be retired from a machine once this module is assigned to it,
which is the first time that has been true. `git` over HTTPS to the mesh's forge starts working,
so the ssh-only clone URL stops being a rule. A module on any machine can call another module's
internal name and verify it.
What got harder: one more module to assign to a machine that needs it, and the machine's trust
store now changes when an assignment changes — which is the point, and is also a thing an operator
can be surprised by. The fetch without prior trust is the same exposure ADR 0098 accepted, now on
every machine that holds the module rather than only where a proxy runs: anything that can stand
in the middle of the mesh's own network at the moment of the fetch can be believed. The mesh
already treats that network as the thing it authenticates.
## References
- [issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md) —
the symptom and the evidence.
- [ADR 0098](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md) — a fact made at
first start is fetched from its provider; this extends it from one program to the machine.
- [ADR 0082](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md) — the precedent
this deliberately does not follow, and why it is right where it is.
- [ADR 0005](0005-the-node-host.md) — the host is where one operating system's difference lives.
- [`03-DESIGN/01-to-be/08-connectivity.md`](../03-DESIGN/01-to-be/08-connectivity.md).
@@ -0,0 +1,166 @@
---
topic: the tiers
status: accepted
date: 2026-09-30
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0066-public-routing-is-name-agnostic.md
---
# 148. The mesh's names are resolved, not copied into every container
## Context
The mesh gives every container it declares the whole roster of mesh names as entries written into
the container's own hosts file at creation
([design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md),
[ADR 0066](0066-public-routing-is-name-agnostic.md)). A container takes those entries once and never
looks again.
Three issues are the same fact arriving three times.
**A container keeps the address it was made with.** Adopting the predecessor's tunnel moved the hub's
private address; the declaration followed it within one push and nothing on the machine did. The
forge's container held the old address, lost its database, reported healthy while its existing
connections lasted, and then the public name went down
([issue 109](../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md)).
**The same fault, four days later, undetected for five days.** One container had restarted 2286 times
against a database it could no longer find, while the mesh reported the machine as doing what it was
told. Inside it, `novox.internal` was an address that had not existed for five days
([issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md)). Forty-eight
other containers were current, none of them corrected — each had been recreated for some other
reason and picked up the roster on the way.
135 was fixed by putting the roster into the digest the host compares a container against, so a
container whose names moved is recreated like one whose image moved. **That made the roster part of
every container's identity**, which is the third arrival:
**One name moving replaces every container in the mesh.** Migrating one small module on one machine
took four routine actions; each changed the roster, and each replaced every container on the control
node — its own store, the registry, the edge proxy, the forge, the directory, mail. The control plane
was unreachable twice while its own store came back through crash recovery
([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). None of
the replaced containers had anything to do with the module being migrated, or with its machine.
The blast radius of a name is now every container that carries the list, which is all of them. The
node-by-node migration ahead adds names one module at a time — on one machine alone that is around
twenty-five — and each would be a full restart of every service on the hub.
## Considered Options
**1. Keep the roster in every container and accept the churn.** Rejected. It is not a cost that can
be paid down: the mesh gets more names as it grows, and every name costs a restart of everything.
A rollback costs another.
**2. Scope each container's entries to the names it actually binds.** A container is given the names
of the things it declared a requirement on, so a name's blast radius is its consumers. Tidy, needs no
new mechanism, and keeps 135's guarantee exactly.
Rejected, and this is the close one. It contradicts the standing intent that **anything on the mesh
can call anything on it** — three cases, same machine, the private network, the public network, and
no fourth. Scoping resolution to declared couplings makes a name reachable only where the mesh was
told in advance that it would be wanted, and a person debugging inside a container would find names
missing that exist everywhere else on the machine. It also leaves the roster in the digest, so the
churn returns the moment a widely-bound name moves — smaller, not gone.
**3. Resolve at lookup time through the machine's resolver, and copy nothing.** Chosen.
## Decision
**A container resolves the mesh's names through its machine's resolver, at the moment it asks. No
mesh name and no mesh address is written into a container, and none is part of a container's
identity.**
The three consequences that make this worth doing:
- **Staleness stops being possible**, rather than being detected. 109 and 135 are not bugs that were
fixed; they are a shape that no longer exists. A name that moves is answered differently by the next
lookup, in every container, with nothing recreated and nothing restarted.
- **A name's blast radius becomes nothing.** Assigning a module on one machine does not touch a
container on another.
- **Anything can still call anything**, which option 2 gave up. The resolver answers every mesh name to
every asker on the machine, exactly as it answers the machine itself.
**The resolver is a machine-level process, not a container** — one of the modules that is not a
container at all — so a container depending on it is not the circularity it would be if the mesh's
own store had to resolve a name through something the store's own runtime had to start first.
**What a module declares for itself is untouched.** Entries a manifest asks for are the module's own,
stay in the container, and stay in its identity: they are part of what the module *is*, they do not
move when the mesh's roster does, and the mesh does not know what they mean.
**The machine's own roster file is untouched.** It is a file, rewritten in place, read by processes and
people; nothing restarts when it changes. It is only the *copy into each container* that this ends.
### The order this lands in, which is not a preference
**Nothing may stop copying names until resolution works from a container.** Removing the copy first
reintroduces 109 and 135 — silently, and on a live mesh, which is exactly how both were found.
1. **A container on any network can reach the resolver.** Today a container on the runtime's default
network asks from an address the converged filter drops, so it has no DNS at all
([issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md));
and on two of four machines the resolver binds loopback only, so the runtime hands containers a
public resolver instead. Both are prerequisites, not related work.
> **Progressive insight — 2026-09-30, later the same day. The loopback claim was wrong.** The
> resolver bound the private address on all four machines; on two the runtime had never been told
> to use it, and on all four the resolver discarded a query that arrived on the runtime's bridge.
> The step stands; the facts under it were those. Both fixed the same day
> ([issue 110's resolution](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/01-resolution.md)),
> and step 3 landed after them.
2. **The runtime is told which resolver to use, per machine, as a file** — not per container as a
creation-time argument, or the resolver's address is back in every container's identity and the
problem has only got smaller.
3. **Then, and only then, the roster leaves the declaration and the digest.**
Until step 3 the mesh keeps copying, and keeps comparing. 151 stays open until step 3 lands; it is not
closed by this record, only answered by it.
## How this is checked
- **A container resolves a name that moved, without being recreated.** Move a name the mesh serves;
from a container that was running before the move and has not been touched since, the name answers
with the new address. This is the one 109 and 135 would both have failed.
- **A name's blast radius is nothing.** Add a routed name on one machine; no container on any other
machine is recreated. The apply report on each machine says nothing changed. This is 151.
- **Anything calls anything.** From a container on any machine, every `<node>.internal` name and every
routed name the mesh serves resolves — including names the module never declared a requirement on,
which is the guarantee option 2 would have given up.
- **On every network the runtime offers.** The first three hold for a container on the runtime's
default network as well as one on a declared network, because the default network is the case that
has no DNS today.
- **No mesh name is in a container's spec.** A test asserts the digest a host computes for a container
does not move when the mesh's roster does, and does move when the module's own declared entries do.
## Consequences
- **The resolver becomes load-bearing for every container**, where before it was load-bearing for the
machine. This is a real cost and is accepted: a resolver that is down is a machine that cannot
resolve, which is already true of the machine itself, and is a smaller event than a roster change
destroying and recreating every container on the machine.
- **Design 08's "a file rather than a resolver" no longer describes containers.** It was written when
the mesh had no resolver and it gave the right answer then. The reasoning it rested on — every Linux
has a hosts file, no package needed — was already overtaken by names a hosts file cannot express:
service names and wildcards under `<node>.internal`, which is why the resolver was built.
- **ADR 0066's mesh-wide propagation is kept and its mechanism changes.** A routed name still reaches
every asker in the mesh; it reaches them through the resolver rather than by being written into each
container. The consequence 0066 records — that an internal issuer's challenge needs the routed name
resolvable inside the mesh — holds unchanged and by the same means the machine already uses.
- **Issue 110 stops being a container-DNS inconvenience and becomes a prerequisite** for the mesh not
restarting itself whenever it learns a name.
- **A container started by hand gets the mesh's names too**, where before only declared containers did.
Design 08 drew that boundary deliberately, on the grounds that reaching into every container is what
a nameserver would be for. This record accepts that consequence rather than working around it: a
person debugging in a hand-started container resolving the same names as everything else is the
behaviour worth having, and it is what "anything can call anything" means.
## References
- [issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md) — one name replaces every container; the question this answers
- [issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md) — the roster put into the digest
- [issue 109](../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md) — the first arrival
- [issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md) — the prerequisite
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — routed names propagate mesh-wide; extended here
- [design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md) — the file-not-resolver reasoning this narrows
@@ -0,0 +1,106 @@
---
topic: building it
status: accepted
date: 2026-09-30
deciders: jochen
reconstructed: false
supersedes: 02-DECISIONS/0068-the-lab-takes-requests.md
---
# 149. The live mesh is the test bed
## Context
[ADR 0068](0068-the-lab-takes-requests.md) proposed that the lab accept queued requests
— a bed and a commit — answer them one at a time from a copy it owns, and expose that through tools so
an agent could start a run and come back to it. It has been `proposed` since 2026-09-12 and nothing was
built.
What happened instead is that the mesh became the thing under test. It runs on four machines; every
fault worth finding in the last month was found on them, and none was found in a bed:
- a container holding an address that had not existed for five days, on the control node
([issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md));
- a machine reading healthy for eleven hours while no module could reach another
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md));
- one name replacing every container on the hub
([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md));
- a consumer assertion that is correct on a mesh being raised and fatal on one that is running
([issue 156](../04-ISSUES/156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md)).
The last is the one that settles it. That change was exercised on the raise path — which is what a bed
*is* — and the raise path is the only path on which the fault cannot appear. A bed raises a mesh; it
does not have a mesh that has been running for weeks, with consumers already bound, containers created
against an older roster, and an adopted machine carrying a predecessor's configuration. **The faults
that cost the most were all faults of a mesh that already exists**, and a bed is by construction a mesh
that does not.
Lab runs are also expensive in a way that changed the behaviour around them: each costs a build and
several minutes, so they were batched, and a batched test is one whose result arrives after the next
three changes were already written.
## Considered Options
**1. Build 0068 as proposed.** Rejected. It answers a question nobody is asking: the bottleneck was
never that a person had to sit at the lab, it was that a bed cannot hold the state the faults live in.
Queueing and tooling a mechanism that finds the wrong class of fault faster is not an improvement.
**2. Leave 0068 `proposed`.** Rejected, and it is why this record exists rather than nothing. A record
that contradicts current practice and sits unresolved is worse than either answer: it reads as intent
to anyone who finds it, and the practice it contradicts is written down nowhere but a handoff note.
**3. Record that the live mesh is the test bed, and supersede 0068.** Chosen.
## Decision
**A change is verified against the mesh that is running.** Not because a bed would be unwelcome, but
because the state that breaks things is state a bed does not have: containers made against an older
roster, consumers already bound, an adopted machine, a store with weeks of history.
**A change that can only be exercised on the raise path is not verified.** If the only test available
raises a fresh mesh, the record says so, and says which case was therefore not covered. The words
"exercised on a fresh mesh" are a statement about coverage, not a pass.
**The lab is not retired**, and [ADR 0016](0016-the-lab.md) stands. It remains the place to raise a
mesh from bare, which is the one thing the live mesh cannot be asked to do and the one thing a bed does
better than anything else. What this record removes is the lab as the *default* answer to "is this
change good", and with it 0068's queue, tools and request protocol.
**Accuracy over a green run.** An honest failure on the live mesh beats a pass in a bed that could not
have failed — and a change that is risky on the running mesh is a reason to make the change smaller,
not a reason to test it somewhere it cannot break.
**What 0068 got right is kept as a rule, not a mechanism:** a run reads a copy that is not anybody's
working tree. Every run of the lab that mattered was pinned to a checkout rather than a worktree, and
the ones that were not produced results about code nobody had written down.
## How this is checked
- **A record that says a change was verified says on what.** Where it was a fresh mesh, it says which
case is uncovered. This is the clause that would have caught 156: its change was verified, honestly,
on the only path where it works.
- **The lab is not in the path of a merge.** No check, playbook or handoff requires a bed to have run.
- **0068 is unreachable as intent.** Its status is `superseded` and it names this record, so a reader
arriving at the queue design finds out immediately that it was not built and why.
## Consequences
- **A fault can be introduced on the machines that serve.** This is the cost, it is real, and it was
paid twice in one evening — a control plane crash-looping for half an hour, and every container on the
hub recreated five times. Both were found in minutes because they were live, and both would have
passed a bed.
- **There is no pre-merge gate beyond the repositories' own suites.** `make check` and the three hq
checks are what stands between a change and the machines, which raises what those suites are worth
and makes a test that cannot fail a genuine defect rather than an untidiness.
- **Raising a mesh from bare is now the lab's whole job**, and is exercised deliberately rather than
as a side effect of testing something else. The foundation work
([issue 146](../04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md))
is that job, and it is also the proof that the mesh can make another of itself.
- **An agent cannot hand a run to a queue and come back**, which 0068 would have given. In practice it
watches a push and reads the machines, which is what happened anyway.
## References
- [ADR 0068](0068-the-lab-takes-requests.md) — superseded by this
- [ADR 0016](0016-the-lab.md) — the lab, which stands
- [issue 156](../04-ISSUES/156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md) — correct on the raise path, fatal on a running mesh
@@ -0,0 +1,113 @@
---
topic: what runs on it
status: accepted
date: 2026-09-30
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md
---
# 150. A module's own code runs as supervised processes under the module's one account
## Context
The repository answers "what runs a module's own code" two ways and reconciles them nowhere
([issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md)).
[ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) is accepted and
says **a container** — "the tool runtime carrying that module's compiled code" — and "one module, one
process, one account". Two `proposed` design documents say a **`process`** resource running an argv,
supervised by the machine, and one of them declares *four* of them for a single module and presents
four as the point. Neither design document names 0047 in its `decisions:`, and the string `process`
as a resource type appears in no decision record at all. The thing as built is the container.
Two things have happened since 0047 was written that bear on it directly.
[**ADR 0142**](0142-the-mesh-delivers-its-own-components-as-binaries.md) decided that the mesh's own
components are binaries on the machine rather than container images, and
[issue 114](../04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md) was closed
by it. That settled the mesh's components and deliberately said nothing about a module's.
And the standing definition of a module hardened: **a module is software that delivers one or more
services, and a module is not a container.** It may deliver them as a container, an installed package
with a unit, a binary, or configuration files; 61 of 73 happen to use a container and 11 do not,
including the resolver, sshd and fail2ban. A rule that a module's *own code* must be a container makes
the one kind of module the mesh writes itself the only kind that has no choice.
## Considered Options
**1. Hold 0047 as written: a container.** Rejected. Its own reasoning does not require one. What 0047
argued for was a runtime **per module** rather than one for the whole node, because a node-wide runtime
could not hold a per-module broker account and per-module runtimes competing on one tool key would each
be handed calls for tools they do not have. A supervised unit per module satisfies that argument
exactly — it is per module, and a unit runs as an account. The container was the mechanism to hand, not
the conclusion.
**2. Let each design document choose.** Rejected; that is the present state and it is what issue 117
reports. A module author reading the guide writes four processes; a module author reading the record
writes a container; nothing tells either that the other exists.
**3. Settle the hosting form as a supervised process, and settle the count separately.** Chosen.
## Decision
**A module's own code runs as one or more supervised processes on the machine, under the module's single
account.** Where [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) says
"a container, the tool runtime carrying that module's compiled code", read this record. Everything else
0047 decided stands untouched: a tool is served on its own key, only the module that serves it answers,
and the module's account is scoped to exactly its tool keys.
**The invariant is the account, not the process count.** 0047's "one module, one process, one account"
carried its weight in the last clause. Its stated worry about a second process was "not a second one to
scope and seal" — a second *identity* to grant, seal a secret to, and scope on the bus. Several
processes sharing the module's one account create no second identity, so nothing further is scoped or
sealed, and a module may therefore declare as many as its work has shapes: events, tools, a
provisioner, a scheduled ingest. **What a module may not have is two accounts.**
**A module that delivers its service as a container still does.** This record is about the code the
module itself carries — its tools, its events, its provisioner — and not about the software it delivers.
A module wrapping a third-party image wraps a third-party image.
**Why supervised by the machine rather than by the mesh:** it is the same answer ADR 0142 gave for the
mesh's own components, for the same reason. A unit the machine restarts needs no image, no registry
pull and no runtime to be up before the mesh's own code can run — which matters most for exactly the
modules whose code the mesh cannot start any other way.
## How this is checked
- **No design document describes a hosting form for a module's own code without citing this record.**
Designs 18 and 20 name it in `decisions:`; this is the gap issue 117's third point reports, and
`cycle.py` already enforces that a to-be design names its decisions.
- **A module declaring several processes resolves to one account.** A test composes a module with more
than one process resource and asserts the mesh mints exactly one broker account for it, scoped to that
module's tool keys and nothing else — which is 0047's invariant stated as an assertion rather than a
sentence.
- **A module's own code does not require the container runtime.** A machine with no container runtime
can still run a module whose code is its own, which is the claim that separates this from option 1 and
is checkable on a machine that has one by asserting the declaration names no image for it.
## Consequences
- **The sidecar port stops being needed.** [ADR 0029](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md)
records that "anything that is a service plus a sidecar currently has to publish a port to talk to
itself", and the host's `network` shape exists partly for it. A process beside the service on the same
machine reaches it without publishing anything, so that pressure goes.
- **Something must supervise, and it is the machine.** This adds a unit per module's code to what the
host writes and owns. The mesh already writes and owns units — `nftables` proves a module can write one
and run it — so the mechanism exists; the count grows.
- **A module's code is delivered, not pulled**, which puts it behind the same gap as the host's own
delivery ([ADR 0141](0141-the-host-delivers-its-own-successor.md), not built): nothing yet delivers a
version of a module's binary to a machine. A container's code arrives by `docker pull`, and this does
not. **This is the cost of the decision and it is not paid**; until delivery exists, a module whose code
is its own is a module somebody places by hand.
- **Issue 117 is answered and its three disagreements close differently:** container-or-unit is decided
here; one-process-or-several is decided here as several under one account; and whether the record was
consulted is fixed by designs 18 and 20 naming this one.
## References
- [issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md) — the contradiction this answers
- [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) — extended; its "a container" clause is settled here
- [ADR 0142](0142-the-mesh-delivers-its-own-components-as-binaries.md) — the same answer for the mesh's own components
- [ADR 0141](0141-the-host-delivers-its-own-successor.md) — the delivery this depends on and which is not built
- [ADR 0029](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md) — the sidecar port this relieves
@@ -0,0 +1,99 @@
---
topic: the tiers
status: accepted
date: 2026-09-30
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0066-public-routing-is-name-agnostic.md
---
# 151. A route's internal name is composed under the node that serves it
## Context
A module that requires a route is given two names from one label: a public one, `<label>.<public
domain>`, and an internal one, `<label>.<node>.internal`
([ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md)). Both were composed
from the node the module runs on.
The two are answered differently. The public name is published into every machine's roster at the
address of the node whose proxy serves it ([ADR 0066](0066-public-routing-is-name-agnostic.md)), so
it reaches the proxy from anywhere in the mesh. The internal name is answered by every machine's
resolver as *anything under a node's name goes to that node*
([design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md)) — the node it was composed from, which is
the consumer's. Where the proxy runs on another machine, that name sends a client to a machine with
nothing listening, while the public name works
([issue 139](../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md)).
Every route on this mesh today is served beside its module, so it has not been seen; `route` is
provided mesh-wide precisely so that stops being true.
Beside it, the roster gave every routed name a second entry with the mesh's suffix appended —
`<name>.<public domain>.internal` — because it composed a full name for every entry as it does for a
machine. That name resolved on every machine, was served by nothing, and was refused by the proxy at
the handshake; the first three names tried while reproducing an unrelated issue were those, and the
evidence pointed at a regression that had not happened
([issue 157](../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md)).
## Considered Options
**1. Keep the consumer's name and publish it at the serving node's address**, as the public name is.
The name stays `<label>.<consumer>.internal` and an exact roster entry overrides the wildcard.
Rejected: it makes `<x>.<node>.internal` mean *goes to that node* except when it does not, which is
the one rule the resolver design states; it needs an entry per route where the wildcard needed none;
and which of an exact entry and a wildcard a resolver answers first is the resolver's business, which
the mesh deliberately does not know.
**2. A proxy on every machine, so the serving node is always the consumer's.** Rejected for this
question: it is a different decision about what `route` is — a node-scoped seat with a mesh-wide
fallback — and this mesh runs one proxy on the hub today. Whatever is decided there, a route served
from another machine must have a name that reaches it.
**3. Compose the internal name under the node that serves the route.** Chosen.
## Decision
**A route's internal name is `<label>.<serving node>.internal` — composed under the node whose proxy
answers the route, which is the machine the request arrives at.** The public name is unchanged:
`<label>.<public domain>` of the node the module runs on, which is where the operator put it.
Where the proxy runs beside the module — every route on this mesh today — the two nodes are one and
nothing changes. Where it does not, the name says where the request goes, which is what a name under
a node's name has always meant.
**A routed name has no mesh form.** The roster publishes it as itself, once, at the serving node's
address. Only a machine has a bare name beside its full one.
What certifies the internal name is unchanged by this: the proxy that terminates it obtains a
certificate from the mesh's authority for the names it is given, and it is given this one.
Taken on the operator's standing instruction to answer the open design questions in the work order.
## How this is checked
- **Composition.** A controller test contributes a route from a module on one node to a proxy offered
from another, gathered the way the controller gathers a consumer's contribution for a provider on
another machine, and asserts the internal name carries the serving node.
- **Publication.** A controller test renders a roster with a machine and a routed name and asserts
the routed name appears as itself, once, and never with the suffix appended.
- **On the mesh.** After the change no machine's roster carries a `<domain>.internal` entry, and a
route's internal name still answers from a container with a certificate from the mesh's authority.
## Consequences
- **A route served from another machine now has a usable internal name.** The first module assigned
that way will resolve, where before it would have resolved to the wrong machine with no error.
- **The internal name of a route can change when its proxy moves.** A route re-homed from one proxy
to another gets a new internal name, as the design's rule implies; clients that dialled the old one
reach the old machine. The public name does not move with the proxy and is the stable one.
- **The roster is one line shorter per routed name**, and a person reading a hosts file no longer
finds names that resolve to a refusal.
- **Issue 139's second question — a per-node route holder — is left open**, and is a decision about
what a seat is rather than about a name.
## References
- [issue 139](../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md) — the question
- [issue 157](../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md) — the alias
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — routed names propagate mesh-wide; extended here
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — how the two names are composed and how far each reaches
- [design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md) — anything under a node's name goes to that node
+61 -6
View File
@@ -48,6 +48,31 @@ form above and dated no earlier than the record's own `date:` — an unmarked ed
violation the reviewer looks for in the diff, and a marked one is legible in the record itself. violation the reviewer looks for in the diff, and a marked one is legible in the record itself.
The git history is the backstop, not the record of intent; the note is the record of intent. The git history is the backstop, not the record of intent; the note is the record of intent.
## A pointer back from what a record changes
A new record naming an old one is not enough. **Where a record changes a mechanism an older record
states — without reversing the decision, so no supersession — the older record gets a dated note
saying where its mechanism now lives.** A reader arrives at the old record by following a citation,
and finds text that is still the decision and no longer the method; nothing in it says a later record
moved the method, and the new record is not in their hands.
> **The mechanism changed — YYYY-MM-DD, by ADR NNNN.** What still stands, what moved,
> and why.
Three examples of the shape, all found by being missed: ADR 0066 still described a routed name being
written into every container after 0148 replaced that with resolution; ADR 0047 still said a module's
code runs in a container after 0150 made it a supervised process; and ADR 0016 still read as though the
lab were the test bed after 0149 said the live mesh is. Each was a citation leading to the wrong
answer, in a record that was not wrong about anything it decided.
**This is not machine-checked, and it cannot be from `extends:` alone.** 102 records extend another and
87 name a parent that does not mention them, which is correct: extending usually means building on a
context, and a one-directional pointer is the right shape for that. What needs a note is the narrower
case where the parent's own text has gone stale, and which case that is, is a judgement — so it is a
rule for the author and the reviewer, and the diff is where it is caught. Making it mechanical would
mean a record declaring the relationship in its frontmatter, which is a change to the record schema and
has not been decided.
The records run in the order the decisions were taken, oldest first. The records run in the order the decisions were taken, oldest first.
**Every decision is a record.** There is no ledger and no index file — if a decision is worth **Every decision is a record.** There is no ledger and no index file — if a decision is worth
@@ -134,6 +159,15 @@ python3 00-META/checks/index.py fail if stale
- **0106** — [The bus is NATS](0106-the-bus-is-nats.md) - **0106** — [The bus is NATS](0106-the-bus-is-nats.md)
- **0116** — [The bus is built in five steps, and the protocol moves with it](0116-the-bus-is-built-in-five-steps.md) - **0116** — [The bus is built in five steps, and the protocol moves with it](0116-the-bus-is-built-in-five-steps.md)
- **0119** — [A taken tunnel's predecessor is retired once the take is proven](0119-a-taken-tunnels-predecessor-is-retired.md) - **0119** — [A taken tunnel's predecessor is retired once the take is proven](0119-a-taken-tunnels-predecessor-is-retired.md)
- **0125** — [The bus is the only broker](0125-the-bus-is-the-only-broker.md) *(superseded)*
- **0127** — [AMQP is a provision, not the bus](0127-amqp-is-a-provision-not-the-bus.md) *(superseded)*
- **0128** — [The mesh bus is required, not ambient](0128-the-mesh-bus-is-required-not-ambient.md)
- **0129** — [A seat carries the protocol of its role](0129-a-seat-carries-the-protocol-of-its-role.md)
- **0130** — [The predecessor is ending, and its broker goes with it](0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md)
- **0131** — [Everything on the mesh speaks to the broker seat, and AMQP is not a provision](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)
- **0132** — [A seat carries the tools its holder must serve](0132-a-seat-carries-the-tools-its-holder-must-serve.md)
- **0134** — [The mesh says what it applied](0134-the-mesh-says-what-it-applied.md)
- **0142** — [The mesh delivers its own components as binaries, not as container images](0142-the-mesh-delivers-its-own-components-as-binaries.md)
### Its tiers, from the bottom up ### Its tiers, from the bottom up
@@ -164,6 +198,9 @@ python3 00-META/checks/index.py fail if stale
- **0098** — [A fact a provider makes at first start is fetched from it, not carried in its manifest](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md) - **0098** — [A fact a provider makes at first start is fetched from it, not carried in its manifest](0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md)
- **0108** — [A route carries the policy applied to a request, and names a secret rather than holding one](0108-a-route-carries-the-policy-applied-to-a-request.md) - **0108** — [A route carries the policy applied to a request, and names a secret rather than holding one](0108-a-route-carries-the-policy-applied-to-a-request.md)
- **0109** — [A package registry seat is one per ecosystem, not one for all of them](0109-a-package-registry-seat-is-one-per-ecosystem.md) - **0109** — [A package registry seat is one per ecosystem, not one for all of them](0109-a-package-registry-seat-is-one-per-ecosystem.md)
- **0126** — [A module declares its own seats; the mesh reserves its own](0126-a-module-declares-its-own-seats.md)
- **0148** — [The mesh's names are resolved, not copied into every container](0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
- **0151** — [A route's internal name is composed under the node that serves it](0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md)
### What runs on them, and how it gets there ### What runs on them, and how it gets there
@@ -196,12 +233,29 @@ python3 00-META/checks/index.py fail if stale
- **0091** — [A mount is declared, and there are three things it can be](0091-a-mount-is-declared-three-ways.md) - **0091** — [A mount is declared, and there are three things it can be](0091-a-mount-is-declared-three-ways.md)
- **0099** — [A step that runs once names what it reads, and runs again when it changed](0099-a-step-that-runs-once-names-what-it-reads.md) - **0099** — [A step that runs once names what it reads, and runs again when it changed](0099-a-step-that-runs-once-names-what-it-reads.md)
- **0110** — [A seat is held by one assignment, from a closed set, and it may deliver a provision](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) - **0110** — [A seat is held by one assignment, from a closed set, and it may deliver a provision](0110-a-seat-is-a-module-assignment-from-a-closed-set.md)
- **0112** — [A module definition names no node, no mesh and no path: everything it needs is a requirement the mesh resolves](0112-a-module-definition-names-no-node-mesh-or-path.md) *(proposed)* - **0112** — [A module definition names no node, no mesh and no path: everything it needs is a requirement the mesh resolves](0112-a-module-definition-names-no-node-mesh-or-path.md)
- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md) *(proposed)* - **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md)
- **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md) *(proposed)* - **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md)
- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md) *(proposed)* - **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md)
- **0117** — [A machine's uplink is a seat: the mesh configures the manager, never the link](0117-a-machines-uplink-is-a-seat.md) - **0117** — [A machine's uplink is a seat: the mesh configures the manager, never the link](0117-a-machines-uplink-is-a-seat.md)
- **0118** — [Undeclaring removes what the mesh made, and gives a unit back the state it was found in](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md) - **0118** — [Undeclaring removes what the mesh made, and gives a unit back the state it was found in](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md)
- **0120** — [A roster fact carries its format as a template: the mesh owns the data, the module owns the format](0120-a-roster-fact-carries-its-format-as-a-template.md)
- **0121** — [A system seat is named for its scope, and a module may define its own](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md)
- **0122** — [A seat is data the controller owns, and a rename is a database update](0122-a-seat-is-data-a-rename-is-a-database-update.md)
- **0133** — [A module owns its migrations, and the mesh owns when they run](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md) *(superseded)*
- **0135** — [A module version prepares its state before it runs](0135-a-module-version-prepares-its-state-before-it-runs.md)
- **0136** — [A step gates its module, not the machine](0136-a-step-gates-its-module-not-the-machine.md)
- **0137** — [A machine says which networks it routes](0137-a-machine-says-which-networks-it-routes.md) *(superseded)*
- **0138** — [An assignment binds an endpoint and says how far it reaches](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md)
- **0139** — [A network is forwarded because a module declared it](0139-a-network-is-forwarded-because-a-module-declared-it.md) *(superseded)*
- **0140** — [The filter constrains what arrives from outside, and says nothing about a machine's own guests](0140-the-filter-constrains-what-arrives-from-outside.md)
- **0141** — [The host delivers its own successor, and versions live side by side](0141-the-host-delivers-its-own-successor.md)
- **0143** — [A consumer verifies the grant it is given](0143-a-consumer-verifies-the-grant-it-is-given.md) *(superseded)*
- **0144** — [Anything on a machine may call anything on it, and that is the whole of "local"](0144-anything-on-a-machine-may-call-anything-on-it.md)
- **0145** — [A module checks what the mesh claims is reachable, and it checks itself](0145-a-module-checks-what-the-mesh-claims-is-reachable.md) *(superseded)*
- **0146** — [Connectivity is checked by name, per hosting form, with a valid certificate](0146-connectivity-is-checked-by-name-per-hosting-form.md)
- **0147** — [A module anchors the mesh's authority on a machine, and takes it away again](0147-a-module-anchors-the-meshs-authority.md)
- **0150** — [A module's own code runs as supervised processes under the module's one account](0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md)
### How it is built ### How it is built
@@ -211,9 +265,9 @@ python3 00-META/checks/index.py fail if stale
- **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md) - **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md)
- **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md) - **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md)
- **0016** — [The lab](0016-the-lab.md) - **0016** — [The lab](0016-the-lab.md)
- **0037** — [Where a module lives](0037-where-a-module-lives.md) *(proposed)* - **0037** — [Where a module lives](0037-where-a-module-lives.md)
- **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md) - **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md)
- **0068** — [The lab takes requests, one at a time, and runs each from its own copy](0068-the-lab-takes-requests.md) *(proposed)* - **0068** — [The lab takes requests, one at a time, and runs each from its own copy](0068-the-lab-takes-requests.md) *(superseded)*
- **0069** — [A module is a repository and a path within it](0069-a-module-is-a-repository-and-a-path.md) - **0069** — [A module is a repository and a path within it](0069-a-module-is-a-repository-and-a-path.md)
- **0076** — [The SDK is a published package, and the toolchain resolves it by version](0076-the-sdk-is-a-published-package.md) - **0076** — [The SDK is a published package, and the toolchain resolves it by version](0076-the-sdk-is-a-published-package.md)
- **0082** — [The registry is reached by name, and the overlay is its security](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md) - **0082** — [The registry is reached by name, and the overlay is its security](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)
@@ -222,6 +276,7 @@ python3 00-META/checks/index.py fail if stale
- **0097** — [A vendor image is a declared build input, and a recipe fetches nothing undeclared](0097-a-vendor-image-is-a-declared-build-input.md) - **0097** — [A vendor image is a declared build input, and a recipe fetches nothing undeclared](0097-a-vendor-image-is-a-declared-build-input.md)
- **0107** — [Persistent data is a directory bind, never a named volume](0107-persistent-data-is-a-directory-bind-never-a-named-volume.md) - **0107** — [Persistent data is a directory bind, never a named volume](0107-persistent-data-is-a-directory-bind-never-a-named-volume.md)
- **0111** — [A build source is on the mesh's git seat, or it is an external repository](0111-a-build-source-is-on-the-git-seat-or-external.md) - **0111** — [A build source is on the mesh's git seat, or it is an external repository](0111-a-build-source-is-on-the-git-seat-or-external.md)
- **0149** — [The live mesh is the test bed](0149-the-live-mesh-is-the-test-bed.md)
### How it is checked ### How it is checked
+44 -1
View File
@@ -2,8 +2,9 @@
layer: to-be layer: to-be
status: in-progress status: in-progress
code: [mesh-host] code: [mesh-host]
updated: 2026-09-22 updated: 2026-09-29
decisions: decisions:
- 02-DECISIONS/0141-the-host-delivers-its-own-successor.md
- 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md - 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
- 02-DECISIONS/0102-the-mesh-writes-into-a-shared-file-never-over-it.md - 02-DECISIONS/0102-the-mesh-writes-into-a-shared-file-never-over-it.md
- 02-DECISIONS/0103-what-an-adopted-node-holds-and-what-its-guard-refuses.md - 02-DECISIONS/0103-what-an-adopted-node-holds-and-what-its-guard-refuses.md
@@ -423,3 +424,45 @@ ignored an instruction and "applied" would be a lie. Applying stays one at a tim
is not. **Checked** by the link's unit tests on the drain, and by the genesis bed's settle wait, is not. **Checked** by the link's unit tests on the drain, and by the genesis bed's settle wait,
which counts on a node catching up to the newest declaration rather than the oldest. which counts on a node catching up to the newest declaration rather than the oldest.
## The host delivers its own successor
*2026-09-29, from a change to the host that could reach no machine —
[issue 142](../../04-ISSUES/142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md),
settled by [ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md).*
A merge builds every changed module and the control plane, and the result reaches the machines running
it with nobody asking. The host was the exception: not a build target, named by no declaration, and
identical on every machine because somebody had copied it there.
The supervision needed for this was already right. A clean exit from the host means it has stood aside,
and the launcher's next turn runs whatever is on disk. Consecutive failed starts are counted, a
rollback happens at the limit, and a second failure halts with the machine named rather than the binary.
What was missing was smaller than it looked: nothing told the running host a successor was waiting, and
the rollback resolved its known-good *version* through one operating system's package manager, which no
machine here used.
Keeping a version rather than a path was the clue. That is only useful to something that can choose
between versions present on the machine — so the versions live side by side:
- **A version arrives as an archive, in a directory named for it.** The ordinary resource, fetched by
digest and checked before anything is unpacked. The path written is never the path being executed, so
replacing a running binary — which the kernel refuses — never comes up.
- **The launcher starts the most recently delivered version**, reading no pointer and following no
link, because the version is in the path.
- **The running host stands aside between reconciles and never inside one.** Standing aside mid-apply is
the half-configured machine this document exists to prevent.
- **A version that completes a reconcile records itself and retires what is older than its
predecessor.** The predecessor stays, because that is what a rollback needs.
- **Rollback starts that predecessor** instead of reinstalling a package: no package manager, no cache
somebody else may clean, and the same script on every operating system.
- **A machine says which host version it runs**, on the report it already sends, so being behind is
answerable at all.
One copy by hand remains, once: the first host that understands versioned directories cannot be fetched
by a host that does not.
*How it is checked* is stated with the decision — a second version delivered to a running machine is
run and reported; the exit follows an in-flight apply rather than interrupting it; a version that will
not start is rolled back once and the second failure halts; a completed reconcile retires what is older
than the predecessor and never the predecessor; and the newest of two delivered versions is the one
that runs.
+22 -1
View File
@@ -11,8 +11,9 @@ code:
- mesh-catalog modules/postgres - mesh-catalog modules/postgres
- mesh-catalog modules/lavinmq - mesh-catalog modules/lavinmq
- mesh-lab test/integration/mesh.test.ts (a bare machine becomes a mesh) - mesh-lab test/integration/mesh.test.ts (a bare machine becomes a mesh)
updated: 2026-09-22 updated: 2026-09-29
decisions: decisions:
- 02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md
- 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md - 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
- 02-DECISIONS/0088-the-foundation-filters-before-anything-listens.md - 02-DECISIONS/0088-the-foundation-filters-before-anything-listens.md
- 02-DECISIONS/0004-a-node-and-how-it-joins.md - 02-DECISIONS/0004-a-node-and-how-it-joins.md
@@ -315,3 +316,23 @@ of a database and pushed to over the broker. What arrived and what did not is th
**One fault, and it was in the joining.** The token did not say what the mesh calls the machine, **One fault, and it was in the joining.** The token did not say what the mesh calls the machine,
so enrolment needed a flag its own help said it did not — and failed at the broker with an empty so enrolment needed a flag its own help said it did not — and failed at the broker with an empty
username. Recorded in ADR 0004 as the fifth thing a token carries. username. Recorded in ADR 0004 as the fifth thing a token carries.
## The mesh's own components arrive as binaries
*2026-09-29 —
[ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md).*
The bundle carries three images and one of them is the controller, *because there is nothing to fetch
it with yet*. That reasoning holds and its conclusion changes: the controller is carried as a **binary**
reference rather than an image reference, pinned by digest exactly as before. Nothing about the bundle's
shape moves — it names a thing and the host fetches it — and the container runtime stops being something
genesis must raise before the control plane can exist. It still raises one, for the store and the broker,
which is where somebody else's software belongs.
The mesh's own components — the host, the controller, the catalogue, the builder, the vault — are
delivered as binaries into directories named for their versions, by the mechanism
[ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) describes. Third-party
software stays a container. The split is not about isolation; it is about who built the thing.
Measured before deciding it: the mesh's own components publish no ports at all, so this opens nothing.
Only the store, the registry and the broker publish, and they are staying as they are.
+157 -3
View File
@@ -7,8 +7,13 @@ code:
- mesh-controller internal/identity/authority.go - mesh-controller internal/identity/authority.go
- mesh-host internal/identity/serving.go - mesh-host internal/identity/serving.go
- mesh-host internal/apply (the service that reflects a rule set) - mesh-host internal/apply (the service that reflects a rule set)
updated: 2026-09-27 updated: 2026-09-30
decisions: decisions:
- 02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md
- 02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md
- 02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md
- 02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md
- 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md
- 02-DECISIONS/0104-a-provision-may-be-answered-by-an-adapter-to-the-predecessor.md - 02-DECISIONS/0104-a-provision-may-be-answered-by-an-adapter-to-the-predecessor.md
- 02-DECISIONS/0106-the-bus-is-nats.md - 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md - 02-DECISIONS/0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md
@@ -286,6 +291,31 @@ hosts file by the runtime. That extends the file decision rather than overturnin
mesh and not chosen by a module: a module that listed the machines would go stale the day one mesh and not chosen by a module: a module that listed the machines would go stale the day one
joins, and a module that did not would be one whose containers cannot reach anything by name. joins, and a module that did not would be one whose containers cannot reach anything by name.
**Superseded for containers by [ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
(2026-09-30).** Everything above still describes the machine's own roster file, which is how it works
and how it will keep working. It no longer describes containers.
Copying the roster into each container made the roster part of each container's identity, so one name
moving replaced every container in the mesh — a module assigned on one machine restarted the store, the
registry, the edge and mail on another
([issue 151](../../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). And it
did not stop the staleness it was meant to: a copy taken at creation is stale the moment the roster
moves, twice found as a container holding an address that had not existed for days
([issues 109](../../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md)
and [135](../../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md)).
**A container resolves the mesh's names through its machine's resolver, at the moment it asks, and
nothing is copied.** The resolver is a machine-level process rather than a container, so nothing
circular is being asked for. It was gated on a container being able to reach the resolver from any of
the runtime's networks
([issue 110](../../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
and landed the day that did, 2026-09-30: the controller writes no mesh name into a container and the
host's digest carries only what the module declared for itself.
The paragraph below states the old boundary, and 0148 deliberately gives it up: a container somebody
started by hand resolves the same names as everything else, because the resolver answers the machine,
not a list of containers.
**The boundary, which is deliberate and worth stating:** *declared* containers. A container **The boundary, which is deliberate and worth stating:** *declared* containers. A container
somebody starts by hand is not the mesh's to configure, and reaching into every container on a somebody starts by hand is not the mesh's to configure, and reaching into every container on a
machine — declared or not — is what a nameserver in `resolv.conf` would be for. machine — declared or not — is what a nameserver in `resolv.conf` would be for.
@@ -296,6 +326,14 @@ machine — declared or not — is what a nameserver in `resolv.conf` would be f
service, the rest is the node — so what resolves is *anything under a node's name*, going to that service, the rest is the node — so what resolves is *anything under a node's name*, going to that
node. What routes it once it arrives is a proxy's, and stays separate. node. What routes it once it arrives is a proxy's, and stays separate.
*2026-09-30.* **So the node in a route's internal name is the one whose proxy answers it**
([ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md)).
Composed from the node the module ran on, the name sent a client to a machine with nothing listening
whenever the proxy ran elsewhere
([issue 139](../../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md));
composed from the serving node, the rule above holds without exception. The public name stays the
module's node's, which is where the operator put it.
**The mesh writes the data and runs no daemon.** One wildcard per machine, from the same set that **The mesh writes the data and runs no daemon.** One wildcard per machine, from the same set that
writes the hosts file. A resolver is third-party software and runs *on* the mesh rather than being writes the hosts file. A resolver is third-party software and runs *on* the mesh rather than being
*of* it: the mesh has no business shipping one, choosing which one, or knowing its configuration *of* it: the mesh has no business shipping one, choosing which one, or knowing its configuration
@@ -371,7 +409,9 @@ can reach from the outside but cannot resolve from the inside is a name it canno
authority of its own. authority of its own.
**So a granted route is published into internal resolution as well** — the routed name to the node **So a granted route is published into internal resolution as well** — the routed name to the node
that serves it, mesh-wide, by the same mechanism that writes the node names. It is *given by the that serves it, mesh-wide, by the same mechanism that writes the node names — and as itself: a routed
name has no mesh form, and the suffixed alias the roster once added beside it resolved to a refusal
([issue 157](../../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md)). It is *given by the
mesh, not chosen by a module*, for the same reason the node names are: a module listing the routes mesh, not chosen by a module*, for the same reason the node names are: a module listing the routes
would go stale the day one changes. The mesh propagates the names it was told to serve and still would go stale the day one changes. The mesh propagates the names it was told to serve and still
knows nothing about what they mean knows nothing about what they mean
@@ -625,6 +665,50 @@ the found firewall reloads and reachable from a container on the node, that a ma
enrols through the openings before and after a reload and a reboot, and that after the flip the enrols through the openings before and after a reload and a reboot, and that after the flip the
declared port is open and the undeclared one closed. declared port is open and the undeclared one closed.
### It filters what arrives from outside, and not what the machine's own guests send
*2026-09-28, preparing the control-node's convergence —
[issue 141](../../04-ISSUES/141-the-forward-chain-does-not-follow-the-modules/00-report.md), settled by
[ADR 0140](../../02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md), which replaces
[ADR 0137](../../02-DECISIONS/0137-a-machine-says-which-networks-it-routes.md) and
[ADR 0139](../../02-DECISIONS/0139-a-network-is-forwarded-because-a-module-declared-it.md).*
Traffic passing through a machine is filtered, because a container's published port is traffic passing
through rather than traffic arriving at the machine itself. Having blocked it, the filter then had to
let the machine's own containers reach outward again — and it did that by listing the address ranges
they sit on. Two of those ranges were constants in this repository's code, and the rest were typed by an
operator after the flip had already cut a workstation's containers off from everything.
**The list was the mistake, not its contents.** A constant describes one machine. A typed range goes
stale and cannot tell a network the mesh made from one a predecessor left behind — on the control-node,
six ranges fall outside the constants and two of the six belong to services the mesh does not run. The
attempt to generate the list from the modules put half the rule set on the machine and made the
derivation advisory. Three records, one list.
**And the mesh has no position on a container reaching outward.** §4 exists to say which port is open
and to whom, which is about what arrives. A container of this machine's own opening a connection
somewhere is not a port opened to anybody, and the addresses it might do that from are the machine's
internal plumbing, which the mesh neither owns nor can know.
So the filter constrains what arrives from **outside** the machine and says nothing about what did not.
Traffic passing through is allowed unless it came in on one of the machine's outward links, and then
only where a declared endpoint's reach admits it (§6). The machine says which of its links face
outside — one fact it reports, like the kind of firewall it found and the tunnel it carries, not a
setting and not a list of addresses. It does not change when a module is added or removed, which is the
whole difference from what it replaces. A machine that has reported no outward link is sent no filter
at all, and keeps the one it has, because a rule written around a link with no name is a rule set that
does not load — a machine filtering nothing while its unit reports success.
Ports go on following the modules exactly as before: assign a module to a machine and the port its
assignment says it reaches on opens. Nothing about a network is said anywhere, by anybody.
*How it is checked:* a bed converges a machine carrying containers on several networks, none of them
named in any setting, and each reaches outward afterwards — which fails against the previous behaviour,
where the same flip cut them off, and is how it was written; a network made *after* the last declaration
needs no new filter; a declared port is reachable from off the private network and an undeclared one is
not; no address of a machine's own networks appears in a rendered filter, asserted on the text; and a
machine reporting no outward link is refused in the control plane with its existing filter left alone.
## 5 — Certificates ## 5 — Certificates
**Two authorities, kept separate on purpose.** **Two authorities, kept separate on purpose.**
@@ -664,6 +748,22 @@ step, so when the authority moves the root is fetched again and the proxy is rec
*How it is checked:* the route-forwarding bed installs the authority, the proxy and a consumer *How it is checked:* the route-forwarding bed installs the authority, the proxy and a consumer
from the catalogue and asserts the routed name is served. from the catalogue and asserts the routed name is served.
**And a machine trusts that authority because a module put its root in its trust store**
([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md)). The proxy's fetch
answers for the proxy and for nothing else: a browser, `git` over HTTPS, a package manager and
every module calling another by an internal name read the machine's own trust store, and the mesh
had never written anything there. A module requiring the authority does the whole of it — fetch
the root over the mesh network, place it where this machine's TLS clients look, refresh the
extracted bundles — and stopping it, which is what being unassigned does, takes the anchor away
and refreshes them again. Not the controller's business, because being on the private network is
what makes the authority *reachable* and is not the same fact as having a reason to *verify* a
mesh name; and because where anchors live and which command refreshes them is one operating
system's difference, which is the host's half of the mesh
([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)).
*How it is checked:* on a machine holding the module a plain client verifies an internal HTTPS
name with no bundle argument, and on one without it the same fetch fails to find an issuer — both
halves, because only the pair tells the anchor apart from something that already trusted it.
### What was built ### What was built
*2026-08-31.* *2026-08-31.*
@@ -710,6 +810,60 @@ that verifies against the internal root and nothing else — which cannot succee
first reached the name to certify it* first reached the name to certify it*
([ADR 0066](../../02-DECISIONS/0066-public-routing-is-name-agnostic.md)). ([ADR 0066](../../02-DECISIONS/0066-public-routing-is-name-agnostic.md)).
## 6 — One statement behind exposure, filtering and certificates
*2026-09-28, preparing the control-node's convergence —
[issue 140](../../04-ISSUES/140-an-endpoints-reach-is-not-declared/00-report.md), settled by
[ADR 0138](../../02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md).*
The three sections above each decide, independently, how far a service reaches. §3 composes a name
from a label and the node's domain. §4 opens a port to the source a listen named. §5 certifies the
names that exist, from whichever authority the proxy holds. Each is coherent on its own, and together
they mean **reachability is never stated anywhere** — it is the sum of three derivations, and a sum is
not something anyone can review or refuse.
What that costs, measured: an identity provider holding a public certificate valid 90 days and an
internal one valid 24 hours, neither asked for by any assignment, because both names existed and a
proxy certifies what it serves. And an endpoint that is not routed — git over ssh — which can be
spoken about only in the filter's vocabulary, so *this must be reachable from outside* is a setting
exactly one mechanism reads.
**An endpoint is the thing that was missing.** A module declares named endpoints: one port it serves,
what it is for, and what it would serve that to absent any instruction. A route contribution names an
endpoint rather than repeating a port number. An assignment — which is where a module's configuration
lives ([ADR 0046](../../02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md))
— then binds each endpoint to a machine port and says how far it reaches.
One value, three readers:
| reach | the filter opens | the proxy serves | the certificate comes from |
|---|---|---|---|
| `internal` | the machine port, to the private network | the internal name | the mesh's own authority |
| `public` | the machine port, to anywhere | the public name | the public authority |
| `both` | the machine port, to anywhere | both names | each name's own authority |
**An endpoint that is not routed is reached and never named.** No route contribution means no name is
composed and no certificate requested, while the filter still acts on it. That is the case the model
could not express at all, and it is the ordinary case for anything that is not HTTP.
**Nothing moves until an assignment says so.** An endpoint whose assignment is silent keeps the
default its manifest states, so every machine already converged renders exactly as it does today —
the same property §4 needed when a machine gained a way to say which networks it routes.
This is what the certificate questions were waiting for. Which authority signs a name, whether a name
may appear in a public issuance log, and what must be trusted where are all answerable once an
endpoint says whether it is internal — and unanswerable while the proxy decides by composing every
name it can. It is also the fact
[issue 139](../../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md)
needs: an internal name should be composed from the machine serving the endpoint, which is the
assignment that bound it.
**How it is checked** is stated with the decision: one module with two endpoints of differing reach
asserted per chain body; the names and the certificate requests following the reach and failing
against today's behaviour, where both are always composed; an unrouted endpoint filtered and never
named; an assignment naming an endpoint the module does not declare refused where it is said; and a
silent assignment rendering byte-identically to today.
## What this removes ## What this removes
The list is worth having in one place, because it is most of the argument: The list is worth having in one place, because it is most of the argument:
@@ -769,7 +923,7 @@ where a found tunnel is left running beside the mesh's; where it is adopted ther
The guard admits the mesh's ports from that one interface, and the predecessor's peers arrive on The guard admits the mesh's ports from that one interface, and the predecessor's peers arrive on
it. it.
*2026-09-27, [ADR 0119](../../02-DECISIONS/0119-a-taken-tunnels-predecessor-is-retired.md).* The *2026-09-27, [ADR 0127](../../02-DECISIONS/0119-a-taken-tunnels-predecessor-is-retired.md).* The
found configuration is kept only until the take is proven — the found unit down, the mesh's found configuration is kept only until the take is proven — the found unit down, the mesh's
interface up and handshaking with a peer. Then it is removed from where the found unit reads it interface up and handshaking with a peer. Then it is removed from where the found unit reads it
(its original stays kept), the hold ends, and the predecessor's tunnel cannot be raised again by (its original stays kept), the hold ends, and the predecessor's tunnel cannot be raised again by
+32 -1
View File
@@ -5,7 +5,7 @@ code:
- mesh-controller internal/builder - mesh-controller internal/builder
- mesh-controller cmd/mesh-controller (build, build --behind, push, status) - mesh-controller cmd/mesh-controller (build, build --behind, push, status)
- mesh-controller internal/inventory/builds.go - mesh-controller internal/inventory/builds.go
updated: 2026-09-21 updated: 2026-09-29
decisions: decisions:
- 02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md - 02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md
- 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md - 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md
@@ -226,3 +226,34 @@ when the current failure began and how many reports in a row have said it — th
id, whatever the words; three make the machine stuck, and `status` says so beside the failure. The host keeps trying — stuck is what the mesh id, whatever the words; three make the machine stuck, and `status` says so beside the failure. The host keeps trying — stuck is what the mesh
knows, not what the machine is told. *How it is checked:* an inventory test counts three identical knows, not what the machine is told. *How it is checked:* an inventory test counts three identical
reports, a different one, and a clean apply; the status test asserts the word appears. reports, a different one, and a clean apply; the status test asserts the word appears.
## Everything may call what is exposed to it, and local is not a boundary
*2026-09-29, from an outage that ran eleven hours —
[issue 145](../../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
settled by [ADR 0144](../../02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md).*
A grant is four facts and a credential: the provision, the machine, the port, and who the consumer is
when it connects. It is the whole mechanism by which anything in the mesh reaches anything else, and it
rests on three things being callable — what runs on the same machine, another machine's service over the
private network where it is exposed there, and another machine's service over the public network where it
is exposed there.
The filter had two of those. A service exposed to the private network admitted the machines' own addresses
on it; a caller on the machine carries such an address, and a caller inside one of that machine's
containers carries a bridge address and matched nothing. Measured, same destination and same machine:
`src 10.10.0.1` from the machine, `src 172.17.0.8` from a container on it. So a module reaching its
database on its own machine's name timed out for eleven hours while the mesh called the machine healthy.
The second case worked by accident: a caller on another machine arrives over the tunnel carrying that
machine's address, which the rule matched. Two of three working is why this read as correct.
**So local is not a boundary this mesh draws, and the filter says so once.** Traffic that did not arrive
from outside the machine and did not arrive over the private network is the machine's own, and is
admitted — for every service there, not per service. Whether the caller is a container, a unit or a shell
decides nothing, because the question is "is this the same machine".
A verification mechanism was drafted for this and withdrawn. It would have reported the outage sooner and
would not have prevented it, and the part of it that was hard — deciding which network position to check
from — existed only because the rule was wrong. Whether the mesh should check that a grant works is still
open, in issue 145; it is not the remedy for a configuration error.
+27 -1
View File
@@ -5,8 +5,10 @@ code:
- mesh-controller cmd/mesh-builder - mesh-controller cmd/mesh-builder
- mesh-controller internal/builder - mesh-controller internal/builder
- mesh-catalog modules/builder - mesh-catalog modules/builder
updated: 2026-09-25 updated: 2026-09-30
decisions: decisions:
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
- 02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md
- 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md - 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md
- 02-DECISIONS/0097-a-vendor-image-is-a-declared-build-input.md - 02-DECISIONS/0097-a-vendor-image-is-a-declared-build-input.md
- 02-DECISIONS/0096-an-upstream-image-is-copied-between-registries.md - 02-DECISIONS/0096-an-upstream-image-is-copied-between-registries.md
@@ -263,3 +265,27 @@ ships one and wrong for code the mesh built, which has no unit until the mesh wr
**Tools, hooks and consumers are not further modes**, which is the test of whether three is the **Tools, hooks and consumers are not further modes**, which is the test of whether three is the
right number: they are loaded by a tool host, and a tool host is a process that stays up. right number: they are loaded by a tool host, and a tool host is a process that stays up.
## The builder compiles the languages the mesh is written in
*2026-09-29 —
[ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md).*
The toolchain list was typescript and python, and only typescript had a base module in the catalogue.
Meanwhile the control plane — written in the language this project is mostly written in — was built as
an image from a hand-written Dockerfile, which is the per-repository incantation this whole mechanism
exists to abolish.
So the list gains Go, with a base module providing the compiler exactly as typescript has one. The
obligation the list's own comment warns about — an SDK carrying the broker client, the event envelope
and tool serving — attaches to a **module** written in a language, not to the language being
compilable. The mesh's own components are not modules in that sense; the host is what applies modules.
**And an artifact says what it targets.** A compiled binary is per operating system, pinned at link
time, and a toolchain deliberately accepts nothing from the module — anything a module could override
there it would be writing a Dockerfile to override. The target is therefore a property of the artifact,
not of the recipe: one artifact declared per target, one build each.
A component's version stops being stamped in at link time. It is unpacked into a directory named for
its version, so it reads its version from its own path, and a build no longer has to know what it will
be called.
+99 -54
View File
@@ -3,12 +3,13 @@ layer: to-be
status: proposed status: proposed
code: code:
- mesh-sdk src - mesh-sdk src
- mesh-tools src/broker-amqp.ts - mesh-tools src/broker-nats.ts (and broker-amqp.ts until the rollout)
- mesh-controller internal/link - mesh-controller internal/link
updated: 2026-09-26 updated: 2026-09-26
decisions: decisions:
- 02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md - 02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md
- 02-DECISIONS/0106-the-bus-is-nats.md - 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md
- 02-DECISIONS/0116-the-bus-is-built-in-five-steps.md - 02-DECISIONS/0116-the-bus-is-built-in-five-steps.md
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md - 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
- 02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md - 02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md
@@ -25,26 +26,17 @@ language and nothing more ([ADR 0074](../../02-DECISIONS/0074-the-wire-is-specif
This is a specification, so it says what is required rather than how anything is arranged. Where it This is a specification, so it says what is required rather than how anything is arranged. Where it
describes current behaviour that is *not yet* specified-and-conformed, it says so. describes current behaviour that is *not yet* specified-and-conformed, it says so.
> **The wire below is the bus being replaced.** *2026-09-26.* > **Rewritten onto NATS, 2026-09-26** (step 3 of
> [ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md) moved the mesh's bus to NATS. Everything > [ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md)). What
> in this document that names an exchange, a queue or a routing key — the event exchanges, the > [ADR 0074](../../02-DECISIONS/0074-the-wire-is-specified-not-the-types.md) decided is untouched:
> durable `<node>.<module>.events` queue, the shared `serve.<key>` queue — describes the transport > a floor plus independent capabilities, an implementation legitimate when it claims less,
> being retired, and the conformance fixtures were captured against it. > identity from the sealed credential, at-least-once with dedup on `x-event-id`, and conformance
> as executable fixtures rather than prose. What changed is the transport beneath all of it —
> exchanges and queues became subjects and streams. The envelope keeps its shape
> ([ADR 0042](../../02-DECISIONS/0042-the-shape-of-an-event-on-the-wire.md)).
> >
> What does **not** change is this document's model, which is the part ADR 0074 decided: a floor > Statements here marked *verified* were checked against a running server while the runtime's
> plus independent capabilities, an SDK that implements what it claims and is legitimate when it > client was written, not reasoned from documentation.
> claims less, identity taken from the sealed credential rather than the environment, at-least-once
> with dedup on `x-event-id`, and conformance as executable fixtures per capability rather than
> prose. The envelope keeps its shape ([ADR 0042](../../02-DECISIONS/0042-the-shape-of-an-event-on-the-wire.md));
> it becomes the message body.
>
> Rewriting the wire sections onto the subjects and streams of
> [design 25](25-the-bus-on-nats.md) §2–§3, and recapturing the fixtures there, is **step 3 of
> [ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md)**. Until that lands, read
> the sections below for what two implementations may not disagree *about*, and design 25 for what
> they will disagree about it *on*. A specification that silently described a retired transport
> would be worse than an absent one, because it reads as current — hence this note rather than a
> quiet edit.
## The shape of it ## The shape of it
@@ -70,22 +62,39 @@ document:
| field | is | required | | field | is | required |
|---|---|---| |---|---|---|
| `url` | an `amqps://` URL carrying the account's user and password | yes | | `url` | a `tls://` URL for the bus, with the account's user and password | yes |
| `fingerprint` | sha256 of the certificate the broker must present | yes for a scoped account | | `fingerprint` | sha256 of the certificate the bus must present | yes for a scoped account |
| `node` | the machine this account was issued for | yes for a scoped account | | `node` | the machine this account was issued for | yes for a scoped account |
| `module` | the module this account was issued for | yes for a scoped account | | `module` | the module this account was issued for | yes for a scoped account |
A plain string rather than a document is a **bootstrap URL** — unscoped, for the moment before a A plain string rather than a document is a **bootstrap URL** — unscoped, for the moment before a
mesh can issue anything. An implementation accepts both and must not treat the second as ordinary. mesh can issue anything. An implementation accepts both and must not treat the second as ordinary.
**`node` and `module` are not decoration: every subject an implementation touches is derived from
them.** Its own namespace is `mesh.mod.<module>`, its consumer is `<node>_<module>`, its inbox is
its own. So a credential without them is refused rather than guessed at — an implementation that
fell back to an environment variable would let anything on the machine decide which module it is,
which is what the identity rule below exists to prevent.
The credential itself is fetched, never carried in a declaration: a declaration is persisted as
state and a sealed secret in a stream is an archive rather than a moment
([design 29](32-what-a-module-declares.md) §10).
### Connecting ### Connecting
- The connection **pins the fingerprint**. It does not trust a certificate authority, and it does - The connection **pins the fingerprint**. It does not trust a certificate authority, and it does
not skip verification. A broker presenting a different certificate is refused, whatever else is not skip verification. A bus presenting a different certificate is refused, whatever else is
true of it. true of it.
- A scoped account **does not declare exchanges**. The foundation owns them; an account that may - **The certificate must also carry a name the bus is dialled by.** *Verified:* the NATS client
declare one is an account that may create a parallel mesh by typo. exposes no hook to replace hostname verification, so pinning no longer makes it redundant the
- An implementation **declares its own queue** and nothing else. way it did on AMQP — the pin happens before dialling and the library's own name check happens
beside it. A certificate without a matching subject-alternative name is refused at connect, by
a library error rather than by anything the mesh says.
- An implementation **creates nothing on the bus**: not a stream, not a consumer, not a subject.
Streams and durable consumers are the controller's alone ([design 25](25-the-bus-on-nats.md)
§3), and a module's account cannot reach the JetStream API to make one. An implementation binds
the consumer the mesh created for it, and if it is absent that is a mesh that has not finished
assigning the module, not something for the module to fix.
### Identity ### Identity
@@ -100,26 +109,49 @@ the credential disagree, the credential wins and the variable is overwritten.
## Capability: events ## Capability: events
### The exchanges ### The subjects
| exchange | carries | | subject | carries |
|---|---| |---|---|
| `mesh.events` | every event | | `mesh.mod.<module>.event.<key>` | an event that module emitted |
| `mesh.events.dead` | what could not be handled | | `mesh.seat.<seat>.event.<verb>` | an event the holder of that role emitted |
### The queue Both are captured by the `EVENTS` stream. **An event's source is enforced rather than claimed**: a
module's account may publish only into its own namespace, so `x-source` cannot disagree with where
the message arrived from.
One **durable** queue per consumer, named `<node>.<module>.events`, with as many bindings as the **The `event` token is load-bearing.** A module's namespace also carries its tool calls
module has patterns. Durable because an event emitted while a module is restarting is exactly the (`mesh.mod.<module>.tool.<tool>`), and a stream is defined by a subject filter — without the token
one that must not be lost. the events stream would capture every tool invocation in the mesh, and a tool call must never be
persisted.
**A message matching two bindings is delivered once**, so an implementation must match the routing ### The consumer
key against its own patterns locally to decide which handlers run. An implementation that ran every
handler whose exchange binding matched would run the wrong one. One **durable consumer** per module, named `<node>_<module>`, carrying one filter per pattern the
module consumes. Durable because an event emitted while a module is restarting is exactly the one
that must not be lost.
**Created by the controller, bound by the implementation.** A module declares what it reacts to
and never how delivery works, so it does not name its consumer, does not choose its ack policy or
delivery limit, and cannot misconfigure them.
*Verified, and it is a trap:* a durable name **may not contain a dot**, while the subject a
consumer acknowledges on is `$JS.ACK.<stream>.<consumer>.…` — two names joined by one. An
implementation that treats them as a single string reads correctly in a permission list and is
refused as a consumer name. Left wrong, the symptom is every message redelivered forever while
the permissions look right.
**One consumer may carry filters wider than one handler's pattern**, because a module subscribing
twice gets one consumer with both. So an implementation still matches the key against its own
patterns locally to decide which handlers run — and **acknowledges a message no handler wanted**,
or it is redelivered until it expires.
### The envelope ### The envelope
Headers ride as AMQP headers. The body is JSON. Headers ride as **NATS headers**; the body is JSON, and the body alone. *Verified:* the payload is
the event's `body`, not the whole envelope re-encoded — an implementation that nested the envelope
would pass every one of its own tests and agree with no other, which is the exact failure the
conformance fixtures exist to catch. The key is recovered from the subject, not carried twice.
| header | is | required | | header | is | required |
|---|---|---| |---|---|---|
@@ -140,6 +172,11 @@ breaking change for everybody.
At-least-once. **Deduplication is on `x-event-id`**, which only the emitter can produce — a At-least-once. **Deduplication is on `x-event-id`**, which only the emitter can produce — a
consumer cannot tell a redelivery from a second event any other way. consumer cannot tell a redelivery from a second event any other way.
On NATS the id does double duty: an implementation passes it as the publish's message id, so the
**server** also refuses a duplicate inside its window. That narrows the window in which a
consumer has to deduplicate; it does not remove the requirement, because the window is finite and
a redelivery after it is still a redelivery.
### What is true, checked (2026-09-16) ### What is true, checked (2026-09-16)
Go emits all five required headers; the SDK requires exactly those. `x-causation-id` and `x-schema` Go emits all five required headers; the SDK requires exactly those. `x-causation-id` and `x-schema`
@@ -154,22 +191,30 @@ version to declare.
A module's tools are its operator-facing surface. A module's tools are its operator-facing surface.
- A tool is served from a **shared durable queue**, `serve.<key>`. Shared, so several runtimes - A tool is served on `mesh.mod.<module>.tool.<tool>`, with a **queue group** — so several
serving one tool compete for a call rather than each answering it. runtimes serving one tool compete for a call rather than each answering it.
- A call is request and reply. The reply returns through the RPC exchange `mesh.rpc`, keyed by the - A call is request and reply on **core NATS, never a stream**. A tool call is not persisted: a
caller's own reply queue — **not** through the default exchange, which would let a caller publish lost one is a timeout the caller already handles, and a stream of them would be the mesh's most
into any queue on the broker. voluminous and least valuable traffic competing for retention with the messages that matter.
- A caller needs a **reply queue**, and that is what a module's scoped account may not declare - The reply goes to the inbox the request carries. A responder may answer it because its account
([issue 049](../../04-ISSUES/049-a-module-can-serve-tools-and-nothing-can-call-them/00-report.md)). is granted **`allow_responses`** — one reply to the subject of a message it actually received,
So a module may serve tools and may not call them. and nothing wider. That is what makes a per-account inbox prefix workable: no user is ever
- **The control plane is the way to ask** granted `_INBOX.>`, so without it a responder could not reach the caller at all.
([ADR 0095](../../02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md)): - **A module may now call a tool, which on AMQP it could not.** *Verified:* two modules on
`ask <module> <tool> [json]` publishes on `mesh.rpc` under `<module>.<tool>` with a private reply separate connections, one serving and one calling, with an answer returned and a throwing
queue bound under its own name, and prints the answer as the module gave it. A module declares handler reaching the caller as an error rather than a timeout.
nothing about being asked — serving a tool is being askable through the control plane. A [Issue 049](../../04-ISSUES/049-a-module-can-serve-tools-and-nothing-can-call-them/00-report.md)
module-to-module call, if one is wanted, is a grant like any other and a later decision. recorded the old limit — a scoped account could not declare the reply queue a caller needs —
*How it is checked:* a tools-only bed asks a served tool through the control plane and asserts and [ADR 0095](../../02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md) routed
an answer arrived, where a timeout would read differently. every ask through the control plane because of it. **That constraint is gone**, and each
account's own inbox prefix replaces it.
ADR 0095 is not thereby reversed: the control plane remains *a* way to ask, and a person asking
a module should still go through it. What changes is that "a module-to-module call, if one is
wanted, is a later decision" is no longer a question about *capability*. It is a policy
question, and the answer the mesh already has is `uses`: a module declares the seat it calls,
and the permission follows the declaration.
- A module declares nothing about being asked — serving a tool is being askable.
--- ---
+2 -1
View File
@@ -5,8 +5,9 @@ code:
- mesh-catalog modules/showcase - mesh-catalog modules/showcase
- mesh-controller internal/builder - mesh-controller internal/builder
- mesh-sdk src - mesh-sdk src
updated: 2026-09-21 updated: 2026-09-30
decisions: decisions:
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
- 02-DECISIONS/0099-a-step-that-runs-once-names-what-it-reads.md - 02-DECISIONS/0099-a-step-that-runs-once-names-what-it-reads.md
- 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md - 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md - 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
+2 -2
View File
@@ -6,7 +6,7 @@ updated: 2026-09-25
decisions: decisions:
- 02-DECISIONS/0084-which-provider-serves-a-consumer.md - 02-DECISIONS/0084-which-provider-serves-a-consumer.md
- 02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md - 02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md
- 02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md - 02-DECISIONS/0126-a-module-declares-its-own-seats.md
--- ---
# 23 — Choosing a provider # 23 — Choosing a provider
@@ -54,7 +54,7 @@ to that provider and not to whichever one is nearest.
**A seat names the mesh's one provider of a kind.** Where a seat delivers the provision, its holder **A seat names the mesh's one provider of a kind.** Where a seat delivers the provision, its holder
answers for it when several providers exist and the consumer named none. That is not picking: the answers for it when several providers exist and the consumer named none. That is not picking: the
choice was made once, mesh-wide, by assigning the holder, rather than once per consumer by naming it choice was made once, mesh-wide, by assigning the holder, rather than once per consumer by naming it
([ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md), ([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) (superseding [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md)),
[26 — The seats](26-the-seats.md)). A named provider still wins over the seat, because a consumer [26 — The seats](26-the-seats.md)). A named provider still wins over the seat, because a consumer
coupled to particular contents has said so. coupled to particular contents has said so.
+170 -32
View File
@@ -1,15 +1,16 @@
--- ---
layer: to-be layer: to-be
status: proposed status: in-progress
code: code:
- mesh-controller internal/link (to be replaced) - mesh-controller internal/link (to be replaced)
- mesh-host internal/link (to be replaced) - mesh-host internal/link (to be replaced)
- mesh-tools src/broker-amqp.ts (to be replaced) - mesh-tools src/broker-amqp.ts (to be replaced)
- mesh-catalog modules/nats (to be written) - mesh-catalog modules/nats (to be written)
- mesh-sdk src (the protocol's NATS binding, step 3) - mesh-sdk src (the protocol's NATS binding, step 3)
updated: 2026-09-26 updated: 2026-09-27
decisions: decisions:
- 02-DECISIONS/0106-the-bus-is-nats.md - 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
- 02-DECISIONS/0116-the-bus-is-built-in-five-steps.md - 02-DECISIONS/0116-the-bus-is-built-in-five-steps.md
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md - 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
- 02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md - 02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md
@@ -17,6 +18,8 @@ decisions:
- 02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md - 02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md
- 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md - 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
- 02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md - 02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md
- 02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md
- 02-DECISIONS/0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md
--- ---
# 25. The bus on NATS # 25. The bus on NATS
@@ -29,19 +32,37 @@ Prose and diagrams only; no configuration is pasted.
## 1. What the bus is for ## 1. What the bus is for
The bus carries five kinds of traffic today, and this design keeps the five, renaming nothing a **The bus is where the mesh happens.** Not a transport the mesh sends things over — the place a
module can see: module is reachable at all, where a role is addressed without knowing who holds it, where the
mesh's own state lives, and where what a module may say is decided by what it declared.
| Traffic | Today | Guarantee it needs | | Traffic | Shape | Guarantee it needs |
|---|---|---| |---|---|---|
| **control** — a node's report, its heartbeat, a build's outcome, an enrolment | queues `control`, `.upgrades`, `.catchup` | nothing lost while the store restarts; retried; in order per node | | **control** — a node's report, a build's outcome, an enrolment | job | nothing lost while the store restarts; retried; in order per node |
| **declarations** — the controller tells a node what to be | queue `node.<name>` | the node gets the newest; a stale one is never applied | | **heartbeat** — a node saying it is alive | fire and forget | none; a lost one is the next one |
| **builds** — the controller asks the build machine to build | queue `builds` | at least once, one builder at a time | | **declarations** — the controller tells a node what to be | state | the node gets the newest; a stale one is never applied |
| **events** — a module says something happened | topic exchange `mesh.events`, keys `<module>.<event>` | delivered to every consumer that declared it; dead-lettered when it cannot be | | **builds** — work for the build machine | job | at least once, one worker at a time |
| **tools** — one module or person asks another's tool a question | exchange `mesh.rpc`, per-tool service queues `serve.<module>.<tool>` | one answer, from one server, or a timeout | | **events** — a module says something happened | 1:many | delivered to every consumer that declared it; dead-lettered when it cannot be |
| **tools** — a module or a person asks another's tool | request/reply | one answer, from one server, or a timeout |
| **work to a role** — a module submits to a capability without knowing who provides it | job | exactly one holder does it; it queues while nobody does |
The sdk's contract — `request`, `handle`, `publish`, `subscribe`, `close` — is the whole surface a The last two rows are the ones worth dwelling on, because they are not messaging in the sense of
module sees, and it does not change ([ADR 0039](../../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)). carrying bytes from A to B. **A role is addressable**, so a caller names the capability and never
the module or the node — and the implementation can be replaced under it without a caller
changing ([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md)). That is a
property of the mesh's architecture that happens to be expressed in subjects.
And more of the mesh lands here as it is built: conditions and observed state in key-value
buckets that anything may watch, the server's own advisories becoming observations like any other
([research 017](../../01-RESEARCH/017-a-mesh-that-heals-itself/00-overview.md)), and a person's
client speaking the bus directly rather than through a surface built over it (§7). None of that
is a message being moved; all of it is the bus being the mesh's centre.
**What a module sees of it is small and derived.** It declares what it emits, consumes, serves
and uses, and the subjects, streams, consumers and permissions all follow from that
([design 29](32-what-a-module-declares.md)). The sdk's contract — `request`, `handle`, `publish`,
`subscribe`, `close` — is the whole surface, and it does not change
([ADR 0039](../../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)).
## 2. Subjects ## 2. Subjects
@@ -52,14 +73,30 @@ permissions are expressed as which branches of it that account may publish to an
mesh.control.<node>.report a node's report (JetStream: CONTROL) mesh.control.<node>.report a node's report (JetStream: CONTROL)
mesh.control.<node>.alive heartbeat (core, no persistence) mesh.control.<node>.alive heartbeat (core, no persistence)
mesh.control.enrol an enrolment request (JetStream: CONTROL) mesh.control.enrol an enrolment request (JetStream: CONTROL)
mesh.control.built a build's outcome (JetStream: CONTROL)
mesh.node.<node>.declare a declaration for a node (JetStream: NODES, last-per-subject) mesh.node.<node>.declare a declaration for a node (JetStream: NODES, last-per-subject)
mesh.build.request work for the build machine (JetStream: BUILDS, work queue) mesh.mod.<module>.event.<event> an event (JetStream: EVENTS)
mesh.events.<module>.<event> an event (JetStream: EVENTS) mesh.mod.<module>.tool.<tool> a tool invocation (core request/reply)
mesh.tools.<module>.<tool> a tool invocation (core request/reply) mesh.seat.<seat>.accept.<verb> work submitted to a role (JetStream: per-seat work queue)
mesh.seat.<seat>.event.<verb> a role's own event (JetStream: EVENTS)
mesh.seat.<seat>.tool.<verb> a role's tool (core request/reply)
mesh.ask.<node>.<command> the controller's command api (core request/reply) mesh.ask.<node>.<command> the controller's command api (core request/reply)
``` ```
**Revised 2026-09-27** ([ADR 0129](../../02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md)):
**`mesh.build.request`, `mesh.control.built` and the BUILDS stream are gone.** A build is work submitted to a role, and the
mesh already has a shape for that — a seat's `accept` subjects, on a work queue with a queue group of
holders, which is what a build queue shared by several machines *is*. Keeping a second mechanism for it
meant two things to reason about and two places for a permission to be wrong. A build's outcome is the
seat's own event — which is why the control branch loses its copy too: one publish reaches whoever
asked, the controller that records it and the catalogue that places it — the fan-out a shared exchange gave for free, as a derived subject rather than a
configured topology.
**Revised 2026-09-26** ([design 29](32-what-a-module-declares.md)): a module's events and tools
moved from `mesh.events.*` / `mesh.tools.*` into one namespace per module, `mesh.mod.<module>.>`,
so a module's authority over its own name is a single subject pattern the server enforces — and
each carries a **kind token**, without which an events stream's filter would capture tool calls.
Seats are the same shape, one namespace per role.
Two things this buys over the exchanges: **request/reply is native** — a tool call is one Two things this buys over the exchanges: **request/reply is native** — a tool call is one
`request` on `mesh.tools.<module>.<tool>` answered by whichever runtime serves it (a queue group per `request` on `mesh.tools.<module>.<tool>` answered by whichever runtime serves it (a queue group per
tool, so several nodes may serve one tool); and **a declaration is last-per-subject** — the NODES tool, so several nodes may serve one tool); and **a declaration is last-per-subject** — the NODES
@@ -69,7 +106,14 @@ exactly the current declaration and nothing older. That is the wire-level answer
*is* the order, and a node that sees sequence n refuses n−1 by construction. *is* the order, and a node that sees sequence n refuses n−1 by construction.
**A reply-to travelling through a JetStream stream is carried in the payload, never in the **A reply-to travelling through a JetStream stream is carried in the payload, never in the
transport `Reply` field.** Revision, first review: core NATS request/reply sets the requester's transport `Reply` field.** *Verified against a running server, 2026-09-27*: a caller published
asking for a reply to `_INBOX.LCr3M83q…`, and the consumer saw a `Reply` field of
`$JS.ACK.PROBE.probe_consumer.1.1.1…`. The address is replaced, not merely at risk — so the
payload-borne reply subject below is necessary rather than defensive, and the check is a test
rather than a note, because a future server that stopped doing this would leave enrolment
working and the reason for the field quietly becoming folklore.
Revision, first review: core NATS request/reply sets the requester's
ephemeral inbox as the message's `Reply` field, and a plain responder answers it directly — but a ephemeral inbox as the message's `Reply` field, and a plain responder answers it directly — but a
message a JetStream consumer delivers has already had that field claimed for the consumer's own message a JetStream consumer delivers has already had that field claimed for the consumer's own
ack address (`$JS.ACK.<stream>.<consumer>...`), so by the time the controller (§3's CONTROL ack address (`$JS.ACK.<stream>.<consumer>...`), so by the time the controller (§3's CONTROL
@@ -89,10 +133,9 @@ Core NATS is at-most-once. Everything the mesh must not lose lives in a JetStrea
| Stream | Subjects | Retention | Why | | Stream | Subjects | Retention | Why |
|---|---|---|---| |---|---|---|---|
| CONTROL | `mesh.control.>` except `alive` | work queue, one consumer (the controller), explicit ack | the store-window guarantee ([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)): the controller `nak`s with a delay while its store is away and the message is redelivered; nothing is dropped | | CONTROL | `mesh.control.>` except `alive` (a build's outcome moved to its seat, ADR 0121) | work queue, one consumer (the controller), explicit ack | the store-window guarantee ([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)): the controller `nak`s with a delay while its store is away and the message is redelivered; nothing is dropped |
| NODES | `mesh.node.>` | last per subject | one declaration per node, always the newest | | NODES | `mesh.node.>` | last per subject | one declaration per node, always the newest |
| BUILDS | `mesh.build.>` | work queue, explicit ack | at least once; a builder that dies mid-build has its message redelivered | | EVENTS | `mesh.mod.*.event.>` | limits (age, size), durable consumer per subscribing module | a subscriber that was down catches up; after `max-deliver` attempts the advisory feeds `mesh.events.dead` (its own small stream) |
| EVENTS | `mesh.events.>` | limits (age, size), durable consumer per subscribing module | a subscriber that was down catches up; after `max-deliver` attempts the advisory feeds `mesh.events.dead` (its own small stream) |
Tool calls and heartbeats stay on core NATS: a lost heartbeat is the next heartbeat; a lost tool Tool calls and heartbeats stay on core NATS: a lost heartbeat is the next heartbeat; a lost tool
call is a timeout the caller already handles. call is a timeout the caller already handles.
@@ -100,6 +143,33 @@ call is a timeout the caller already handles.
Streams and consumers are objects the controller creates at genesis and asserts on start; a module Streams and consumers are objects the controller creates at genesis and asserts on start; a module
declares nothing about them. The controller is the only writer of stream definitions. declares nothing about them. The controller is the only writer of stream definitions.
### The store window, and what moving it into the server changes
The guarantee ([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)) is
that a push the controller cannot record because its store is restarting is **held and retried** —
never dropped, never falsely acknowledged. Here that is a `nak` with a delay: the server holds
the message and redelivers it, so the controller keeps no list of parked messages and one that
restarts mid-window loses nothing it was holding.
**That is a plain win, and it introduces one problem worth naming.** Holding a delivery in memory
let the controller drop an older report when a newer one for the same node arrived, because
acting on the older after the newer would undo the newer. A `nak`ed message belongs to the server
and comes back whatever happened meanwhile — so the older report is redelivered *after* the newer
was applied.
The answer was already in the message. A report carries the **digest of the declaration it is
about**, which exists because an earlier attempt to order reports by time lost the race it
invited: an apply that began under the previous declaration finishes after the next is sent, and
its report reads as newer than the send. Clocks cannot answer *which*.
So supersession stops being something the controller remembers and becomes something it checks —
a report whose digest is not the one outstanding for that node is acknowledged without being
acted on. The same shape as a node refusing a superseded declaration by sequence
([issue 107](../../04-ISSUES/107-a-declaration-carries-no-order/00-report.md)): **ordering settled
by what a message says, not by when it arrived.** And staleness is checked before the store is
waited on, so a redelivery that lost its race does not hold a slot in the window that a current
message needs.
## 4. Accounts ## 4. Accounts
[ADR 0043](../../02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) says a [ADR 0043](../../02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md) says a
@@ -128,19 +198,56 @@ expresses this exactly, per subject, and better than a vhost could:
permissions name only that one prefix, for the reply to any request it makes and nothing permissions name only that one prefix, for the reply to any request it makes and nothing
wider. First review found the account note without this and read it as "any user may wider. First review found the account note without this and read it as "any user may
subscribe any inbox" — which was accurate against the text as it stood. subscribe any inbox" — which was accurate against the text as it stood.
**And a scoped inbox needs `allow_responses`, or nothing can answer.** Revision, found while
composing the first real configuration: the rule above scopes each user's inbox to itself,
which is right — and leaves a responder unable to reply, because the answer goes to the
*caller's* inbox, which the responder has no permission for. The two ways out are granting
every responder `_INBOX.>`, which is exactly the blanket grant this bullet refuses, or NATS's
own `allow_responses`: the server permits one reply to the reply-subject of a message the user
actually received, within a TTL, and nothing else. So authority to answer is bounded by having
been asked, and only principals that serve something are granted it — a pure consumer gets
nothing. Without this the scoping is not merely incomplete: every tool call in the mesh times
out, and the permission list looks correct while it happens.
- **One user per module per node**, as today, with publish permissions - **One user per module per node**, as today, with publish permissions
`mesh.events.<module>.<event>` for each emit, `mesh.tools.<module>.>` to serve its tools, its `mesh.events.<module>.<event>` for each emit, `mesh.tools.<module>.>` to serve its tools, its
own ack-reply subject for each durable consumer it holds, and its own inbox prefix; subscribe own ack-reply subject for each durable consumer it holds, and its own inbox prefix; subscribe
permissions for each consumed event's subject, its tool subjects, and that same inbox prefix. permissions for each consumed event's subject, its tool subjects, and that same inbox prefix.
Nothing else. A module that tries to publish outside its emits is refused by the server, not by Nothing else. A module that tries to publish outside its emits is refused by the server, not by
convention. convention.
- **The controller's user** owns `mesh.control.>`, `mesh.node.>`, `mesh.build.>` and the streams. - **The controller's user** owns `mesh.control.>`, `mesh.node.>` and the streams, and may submit work
to the seats the mesh's own flows use — a build, for one (ADR 0121).
**A host's user** may publish its own `mesh.control.<node>.>` and subscribe its own **A host's user** may publish its own `mesh.control.<node>.>` and subscribe its own
`mesh.node.<node>.declare` — and nothing of any other node's. `mesh.node.<node>.declare` — and nothing of any other node's.
- **A person's user** (§7) is a module-shaped user with permissions on the tool subjects it may - **A person's user** (§7) is a module-shaped user with permissions on the tool subjects it may
invoke, issued and revoked by the controller like any account. invoke, issued and revoked by the controller like any account.
**Accounts are configuration, not API calls.** The controller composes the server's user list and **The server does not verify client certificates, and TLS is still required.** *Revision,
2026-09-27, found by building the module's image and connecting to it as a host would.* The first
composed configuration said `verify: true`, which makes the server demand a **client** certificate —
and nothing in the mesh presents one. A host pins this server's exact certificate and authenticates
with the password the mesh minted ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)),
and a module's runtime does the same. With it on, every connection in the mesh dies at the TLS
handshake before any password is looked at, and the error — "client didn't provide a certificate" —
reads as a fault in the client rather than in the bus's configuration. The `tls` block is what makes
TLS required; `verify` only decides whether client certificates are checked. What is given up is a
second factor the mesh has no machinery to issue or rotate — a certificate per module per node — and
what is kept is stronger than a name check in both directions: an exact pin outward, a password
scoped per user inward. **Mutual TLS is a later question and would need that machinery first.**
**The mesh composes the accounts; the module composes its server.** *Revision, 2026-09-27, while
building the composition.* An earlier reading of the paragraph below had the controller writing the
whole file. It writes only the user list. A server's ports, its TLS paths and its store directory
are properties of the container the module raises — they live in its image and its mounts and change
when it does — so the module declares its own configuration and `include`s the mesh's half. A
controller that wrote the whole file would have to be kept in step with a Dockerfile it never sees,
and a module could not change its own image without the mesh agreeing. Asking for the user list is
not enough to receive it: the file holds every user's password hash, so the claim on `mesh-broker` is
what authorises it. And the two files share one directory of necessity — an absolute include path is
resolved relative to the including file's own directory, so a server given one from elsewhere looks
for it underneath that directory and refuses to start.
**Accounts are configuration, not API calls.** The controller composes the mesh's user list and its
permissions into a file the host declares. **How that file reaches the running server is §5's, permissions into a file the host declares. **How that file reaches the running server is §5's,
not this one's** — revision, first review: an earlier draft said "reloads" and cited a precedent not this one's** — revision, first review: an earlier draft said "reloads" and cited a precedent
that does not apply to a container (see §5). No management API, no credential travelling through a that does not apply to a container (see §5). No management API, no credential travelling through a
@@ -157,9 +264,14 @@ signing hierarchy for nothing.
`nats` is a catalogue module claiming the seat `mesh-broker` `nats` is a catalogue module claiming the seat `mesh-broker`
([ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md): the seat ([ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md): the seat
is the server, and the server changes). It declares one container (a single binary; JetStream on a is the server, and the server changes). It declares one container (a single binary; **JetStream on a host
named volume), its listening ports — client, TLS, and the monitoring endpoint on loopback — and a directory bind, not a named volume** — revision, second review:
configuration file the controller composes (accounts, permissions, TLS, JetStream). [issue 115](../../04-ISSUES/115-a-named-docker-volume-is-invisible-and-one-flag-from-gone/00-report.md)
is resolved, and converted the store, the broker and two others away from named volumes for the
reason it names; the bus's own data is not the place to reintroduce one), its listening ports —
**the client port, which carries TLS itself** rather than standing beside a plaintext one as the
AMQP broker's 5671/5672 pair did, and the monitoring endpoint on loopback — and a configuration
file the controller composes (accounts, permissions, TLS, JetStream).
**How that file's changes reach the running server, corrected on revision.** First review: the **How that file's changes reach the running server, corrected on revision.** First review: the
earlier draft named `reload-on` as the mechanism, citing the container runtime's own trust file as earlier draft named `reload-on` as the mechanism, citing the container runtime's own trust file as
@@ -186,10 +298,33 @@ same place `modules/gitea/token.ts` keeps its own state rather than asking the h
The host's only job is what it already does for any directory resource: keep the file's content The host's only job is what it already does for any directory resource: keep the file's content
current. Nothing is declared as `reload-on` or `restart-on` for this resource at all. current. Nothing is declared as `reload-on` or `restart-on` for this resource at all.
Its guard is the same rule as the AMQP broker's: the monitoring port is refused from anything but Its guard is the same rule as the deprecated broker's: the monitoring port is refused from anything but
the private network. It is raised at genesis like the store, adopted as a module in the same the private network. It is raised at genesis like the store, adopted as a module in the same
phase. The predecessor's AMQP broker remains a module of its own, `lavinmq-compat`, with a single phase.
purpose and a retirement condition: no client connected for a period the operator sets.
**The deprecated broker is an ordinary module, not a compatibility layer.** Revision, second review
([ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md): AMQP is not a provision, and everything speaks to the `mesh-broker` seat)): earlier text here, and
ADR 0106 before it, called it `lavinmq-compat` — one purpose, the predecessor's clients, and a
retirement condition of no client connected for a period the operator sets. It is none of those.
A module may legitimately need an AMQP broker as a **backing service**, the way it needs a
database, and the provider that answers that is an ordinary module like any other: no seat, not
foundation, never raised at genesis, installed when something wants it and absent from a mesh
that does not. There is no retirement condition, because the day its last client disappears is
not a day anything is waiting for.
**Revised 2026-09-27** ([ADR 0130](../../02-DECISIONS/0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md)):
**that day is coming.** The predecessor is deprecated — some of it still running, none of it being
migrated, left to stop rather than moved — so the broker retires once nothing requires `amqp`. Still no
retirement *condition* and no end-date machinery: a provision with no consumers has its provider
unassigned, which is the ordinary mechanism and is ADR 0127 being paid off rather than revised. What
also goes with it is the tooling that reaches this installation's machines remotely, because the
predecessor's own mesh talks over that broker — so the rollout is driven from the node, or before the
broker stops.
What is deprecated is AMQP as **the mesh's transport**, which is this whole document. The rule
that remains is about direction rather than software: *inter-module communication goes over the
bus.* A module may hold a broker, a database or a cache for itself; it may not use one as a
channel to another module.
## 6. Joining: the enrolment handshake ## 6. Joining: the enrolment handshake
@@ -272,7 +407,7 @@ wrong until there is a second mesh. *Ends at: the genesis bed.*
**Step 2 — adoption puts the broker in its seat.** A mesh already running does not get a foundation **Step 2 — adoption puts the broker in its seat.** A mesh already running does not get a foundation
module by being raised again; it adopts one in place module by being raised again; it adopts one in place
([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)). The ([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)). The
server is raised beside the AMQP broker on its own ports, carrying no mesh traffic yet, and the server is raised beside the deprecated broker on its own ports, carrying no mesh traffic yet, and the
`nats` module is adopted onto it. The seat it claims is **`mesh-broker`**, unchanged — the `nats` module is adopted onto it. The seat it claims is **`mesh-broker`**, unchanged — the
foundation seats are named after the server's role rather than the product foundation seats are named after the server's role rather than the product
([ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md)) for ([ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md)) for
@@ -333,7 +468,7 @@ itself enforces and a plain client can therefore check:
watcher-poll interval, without a restart. watcher-poll interval, without a restart.
**Step 2 — the adoption bed**, a mesh already running that has never had this server: **Step 2 — the adoption bed**, a mesh already running that has never had this server:
- the server is raised beside the AMQP broker on its own ports and the `nats` module is adopted - the server is raised beside the deprecated broker on its own ports and the `nats` module is adopted
onto it in place, holding the data and the configuration it was raised with; onto it in place, holding the data and the configuration it was raised with;
- the seat it claims is `mesh-broker`, and a second assignment of it anywhere in the mesh is - the seat it claims is `mesh-broker`, and a second assignment of it anywhere in the mesh is
refused at resolution — *one per mesh*, as ADR 0079 requires; refused at resolution — *one per mesh*, as ADR 0079 requires;
@@ -398,8 +533,11 @@ find what changed and why.
**Still open:** **Still open:**
- Whether EVENTS should be one stream or one per emitting module (retention per module vs. one - ~~Whether EVENTS should be one stream or one per emitting module.~~ **Closed**
policy). One stream is proposed; the review may disagree. ([ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md): AMQP is not a provision, and everything speaks to the `mesh-broker` seat) (superseding [ADR 0125](../../02-DECISIONS/0125-the-bus-is-the-only-broker.md))): one stream, and not as a
preference — the streams the bus is made of are composed as configuration before any module
runs, because a provisioner is itself a module that needs a bus account to start. Bootstrapping
decides it.
- The heartbeat interval and the controller's "quiet" threshold on core NATS without persistence — - The heartbeat interval and the controller's "quiet" threshold on core NATS without persistence —
the same numbers as today are proposed. the same numbers as today are proposed.
- Whether the person's client is a catalogue module (runs on an enrolled workstation node) or a - Whether the person's client is a catalogue module (runs on an enrolled workstation node) or a
+103 -30
View File
@@ -4,13 +4,17 @@ status: implemented
code: code:
- mesh-controller internal/catalogue/seats.go - mesh-controller internal/catalogue/seats.go
- mesh-controller internal/catalogue/resolve.go - mesh-controller internal/catalogue/resolve.go
- mesh-controller internal/inventory/seats.go
- mesh-controller internal/inventory/migrations/0039-a-seat-is-held-by-one-assignment-on-record.sql
- mesh-controller cmd/mesh-controller/seats.go - mesh-controller cmd/mesh-controller/seats.go
- mesh-controller cmd/mesh-controller/source.go - mesh-controller cmd/mesh-controller/source.go
- mesh-controller internal/inventory/migrations/0032-a-source-may-live-on-a-seat.sql - mesh-controller internal/inventory/migrations/0032-a-source-may-live-on-a-seat.sql
- mesh-catalog modules/gitea/module.json - mesh-catalog modules/gitea/module.json
updated: 2026-09-27 updated: 2026-09-27
decisions: decisions:
- 02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md - 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
- 02-DECISIONS/0126-a-module-declares-its-own-seats.md
- 02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md
- 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md - 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md
- 02-DECISIONS/0109-a-package-registry-seat-is-one-per-ecosystem.md - 02-DECISIONS/0109-a-package-registry-seat-is-one-per-ecosystem.md
- 02-DECISIONS/0117-a-machines-uplink-is-a-seat.md - 02-DECISIONS/0117-a-machines-uplink-is-a-seat.md
@@ -41,6 +45,31 @@ is refused. A seat makes a role singular, never a module.
**The seat points at the assignment.** Everything the mesh knows about the holder is what it knows **The seat points at the assignment.** Everything the mesh knows about the holder is what it knows
about that assignment: the node, the node's settings for the module, and what the module serves. about that assignment: the node, the node's settings for the module, and what the module serves.
**Which assignment holds a seat is a fact on record, and changes as one act.** Revision, 2026-09-27
([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)). Until
then the holder was derived — the module that is assigned and claims the seat holds it, and a second
eligible assignment was refused. That has no way to pass a seat from one holder to the next without a
moment in which nobody holds it, and the controller finds its own bus through one of these seats: that
moment took the control plane down for an evening. So the holder is now one row the controller keeps,
written by a handover — `seat <name> --to <node>/<module>` — that names the seat and the assignment
taking it over and replaces the previous holder in the same write. Between two handovers the seat has
exactly one holder, and it is never none.
Three consequences follow. **A seat with no row is held as it always was**: the sole eligible
assignment holds it, and two eligible ones are refused — so a mesh that has never handed a seat over
behaves exactly as before, and the row appears the first time somebody does. **With a row, any other
assignment whose module could hold the seat is eligible and silent**: neither refused nor holding.
That is what lets the next holder run beside the current one until the handover, which the bus's move
needs ([28](28-building-the-bus.md), task 5.3). **And a holding is the assignment's**: unassigning the
holder takes the row with it, so a seat never points at something that is not running anywhere, and
the seat falls back to derivation rather than to nothing.
The handover refuses what would make the new holder wrong before anything is written: the seat must
exist, the assignment must exist, and the module must be able to hold the seat — claim it at its scope
and provide what it delivers, judged against the store's row and not against anything compiled into a
binary. It does not check that the module is running yet; `push` confirms that afterwards, and a
handover that could only be recorded after the new holder was up could not be the switch.
**The set is closed.** A seat the mesh does not define is refused wherever it is named, and so is one **The set is closed.** A seat the mesh does not define is refused wherever it is named, and so is one
named at the wrong scope. Adding a seat is a decision, recorded, for the reason every addition to the named at the wrong scope. Adding a seat is a decision, recorded, for the reason every addition to the
host's vocabulary is one: the set is what a person reads to learn what a mesh can have, and an entry host's vocabulary is one: the set is what a person reads to learn what a mesh can have, and an entry
@@ -48,28 +77,55 @@ nobody argued for is an entry nobody can explain.
## The set ## The set
| seat | scope | delivers | typically held by | **The set is derived, and only the mesh's half is written here.** Revision, 2026-09-26
|---|---|---|---| ([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md), superseding
| `mesh-controller` | mesh | — | the controller | [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md)): a module
| `mesh-store` | mesh | — | the store the mesh's own records live in | declares its own seats with their protocols, so the seats a mesh has are the mesh's own **plus
| `mesh-broker` | mesh | — | the broker carrying the mesh's own bus | every registered module's**. The set is still closed — a seat named nowhere is refused — but it is
| `mesh-vault` | mesh | `secret`, reserved | the vault | computed from the catalogue rather than maintained by hand, which is the property 0110 actually
| `the-artifact-store` | mesh | `artifact-store` | the artifact registry | needed and the table could not keep.
| `the-catalogue` | mesh | — | the catalogue |
| `npm-package-registry` | mesh | `npm-package-registry` | the forge |
| `git` | mesh | `git` | the forge |
| `the-build-machine` | node | — | a builder |
| `the-dns-port` | node | — | the local resolver |
| `the-intrusion-prevention` | node | — | an intrusion-prevention service |
| `the-packet-filter` | node | — | the packet filter |
| `the-private-network` | node | — | the private network the mesh runs over |
| `the-resolver-configuration` | node | — | whichever of the alternative resolver configurations is chosen |
| `the-showcase` | node | — | the showcase module |
| `the-uplink` | node | — | the program that manages the machine's own network ([ADR 0117](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md)) |
The controller holds this set in code, and a test asserts both its size and that every entry names **And the mesh's own half is data, named for its scope.** Revision, 2026-09-27, reconciling two
the record that made it a seat. **This table and [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md) records made in parallel: [ADR 0121](../../02-DECISIONS/0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md)
govern, and code that disagrees is what is wrong.** The implementation in progress predates several names a system seat for the scope it is held at — `mesh-*` for one per mesh, `node-*` for one per
machine — and [ADR 0122](../../02-DECISIONS/0122-a-seat-is-data-a-rename-is-a-database-update.md) moves
the set out of compiled code into a table the controller owns, so a rename is one write rather than a
rebuild of everything that names one.
So the set has two halves and neither is written out here: the mesh's own, which the controller holds
as rows, and every registered module's, which is computed from the catalogue. What this document keeps
is what a seat *is* — the rest would be a third copy, stale the first time somebody renamed one, which
is the fault ADR 0122 exists about.
**Every seat below is named `mesh-*`, and the prefix is the reservation rule**: a module declaring
any `mesh-*` name is refused at registration, so there is no reserved-names list to drift. Ten of
these are renamed to restore [ADR 0079](../../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md)'s
convention, which later seats departed from.
| seat | was | scope | delivers | typically held by |
|---|---|---|---|
| `mesh-controller` | — | mesh | — | the controller |
| `mesh-store` | — | mesh | — | the store the mesh's own records live in |
| `mesh-broker` | — | mesh | `mesh-bus` | the broker carrying the mesh's own bus |
| `mesh-vault` | — | mesh | `secret`, reserved | the vault |
| `mesh-artifact-store` | `the-artifact-store` | mesh | `artifact-store` | the artifact registry |
| `mesh-catalog` | `the-catalogue` | mesh | — | the catalogue |
| `mesh-npm-package-registry` | `npm-package-registry` | mesh | `npm-package-registry` | the forge |
| `mesh-git` | `git` | mesh | `git` | the forge |
| `mesh-build-machine` | `the-build-machine` | node | — | a builder |
| `mesh-dns-port` | `the-dns-port` | node | — | the local resolver |
| `mesh-intrusion-prevention` | `the-intrusion-prevention` | node | — | an intrusion-prevention service |
| `mesh-packet-filter` | `the-packet-filter` | node | — | the packet filter |
| `mesh-private-network` | `the-private-network` | node | — | the private network the mesh runs over |
| `mesh-resolver-configuration` | `the-resolver-configuration` | node | — | whichever of the alternative resolver configurations is chosen |
| `mesh-showcase` | `the-showcase` | node | — | the showcase module |
The controller holds **the mesh's own** entries in code, and a test asserts their size and that
every one names the record that made it a seat. A module's seats are not here and never will be —
they are read from the catalogue. **This table and
[ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) govern, and code that
disagrees is what is wrong.** The implementation in progress predates several
things here: seats held by assignments rather than claimed by definitions, the `mesh-vault` seat and things here: seats held by assignments rather than claimed by definitions, the `mesh-vault` seat and
its reservation, and the foundation's seats delivering nothing. It is brought to this table before it its reservation, and the foundation's seats delivering nothing. It is brought to this table before it
merges. merges.
@@ -121,6 +177,15 @@ moves that to a host port requirement.
## A seat that delivers nothing ## A seat that delivers nothing
**A module's declared seat may promise nothing too, and that is a marker seat.** Correction of fact,
2026-09-27: [ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md)'s "a declared seat
carries a protocol" governs what a *holder* must satisfy, not that every seat offers something. A
marker seat's protocol is satisfied by holding it, which is the whole of what a marker says. Refusing
an empty one would refuse most of the node-scoped set, the showcase module's own seat included.
Nothing reaches that state by accident: an unknown manifest field is refused outright, so an empty
protocol was written as one. Checked by a registration test accepting a node seat with no protocol
and by the showcase manifest, which declares one.
Most node seats deliver nothing. They say which module is this machine's packet filter, or which of Most node seats deliver nothing. They say which module is this machine's packet filter, or which of
two alternative resolver configurations it runs, and a second holder is refused. That is the whole of two alternative resolver configurations it runs, and a second holder is refused. That is the whole of
their job, and it is a real one: it is the mesh saying what a machine is, in words a person can read. their job, and it is a real one: it is the mesh saying what a machine is, in words a person can read.
@@ -157,16 +222,24 @@ still to take.
## How it is checked ## How it is checked
The rules here are [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md)'s The rules here are [ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md)'s —
and [ADR 0111](../../02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md)'s, and each is which supersedes [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md)
and keeps every rule below except how the set is formed — and
[ADR 0111](../../02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md)'s. Each is
checked as their tables say: checked as their tables say:
| Rule | Checked by | | Rule | Checked by |
|---|---| |---|---|
| The set is closed, and every entry names its decision | 0110: a unit test on the set's size and decisions; manifest tests refusing an unknown seat or the wrong scope. | | The set is closed, and every mesh entry names its decision | 0118: a unit test on the mesh's own entries. **The refusal of an unknown seat is at registration, not in the parser** — correction of fact, 2026-09-27: a module may hold a seat *another* module declares, which is the point of naming the seat and not its provider, so whether a claimed name exists is a fact about the whole catalogue and a manifest in isolation cannot be judged on it. Registration tests cover an invented name and a name another module declares; the parser still refuses a claim on the mesh's own `mesh-*`/`node-*` namespace and a scope that disagrees with a declaration in the same manifest. |
| A seat is held by one assignment, and only by one whose module can hold it | 0110: resolution tests for a second holder and for a seat the definition does not name. | | The set is derived, and enumerating it is a query | 0118: the overview lists the mesh's own plus every registered module's, asserted against a fixture mesh. |
| A requirement naming a seat is answered by its holder; a foundation seat cannot be named | 0110: resolution tests with a second provider on the consumer's node, with the seat unheld, and naming `mesh-store`. | | `mesh-*` is the mesh's, and a module may not declare one | 0118: a registration test refusing a manifest that declares any `mesh-*` seat, naming the prefix. |
| Several providers and none local is a person's choice | 0110: an assignment test listing candidates with the seat's holder first and recording the pin. | | Two modules cannot declare the same seat | 0118: a registration test; the second is refused and the first untouched. |
| `secret` is reserved | 0110: the parser and resolution refusals for another provider and a pin. | | A holder satisfies the seat's protocol | 0118: a claim whose module does not serve what the seat declares is refused at assignment. |
| Holdings are derived, and the overview lists every seat | 0110: the `seats` command test, including an unheld seat. | | A seat is held by one assignment, and only by one whose module can hold it | 0118: resolution tests for a second holder and for a seat the definition does not name. 0131: `CanHold` is the one judgement, shared by registration and the handover, and its test follows the store's row. |
| A holder on record settles the seat; another eligible assignment is silent, not refused | 0131: resolution tests with a recorded holder on the same machine, on another machine, and under a seat's former name; without a record, the old rule's tests still pass unchanged. |
| A handover replaces the holder as one write, needs an assignment to point at, and goes with it | 0131: store tests — a second handover leaves one row; a handover to a module not assigned where named is refused; unassigning the holder removes the row. |
| A requirement naming a seat is answered by its holder; a foundation seat cannot be named | 0118: resolution tests with a second provider on the consumer's node, with the seat unheld, and naming `mesh-store`. |
| Several providers and none local is a person's choice | 0118: an assignment test listing candidates with the seat's holder first and recording the pin. |
| `secret` is reserved | 0118: the parser and resolution refusals for another provider and a pin. |
| Holdings are derived, and the overview lists every seat | 0118: the `seats` command test, including an unheld seat. |
| A build source on the seat records no address; an unheld seat refuses only self-hosted builds | 0111's tests. | | A build source on the seat records no address; an unheld seat refuses only self-hosted builds | 0111's tests. |
@@ -8,7 +8,7 @@ decisions:
- 02-DECISIONS/0115-one-assignment-of-a-module-per-node.md - 02-DECISIONS/0115-one-assignment-of-a-module-per-node.md
- 02-DECISIONS/0113-the-vault-makes-every-secret.md - 02-DECISIONS/0113-the-vault-makes-every-secret.md
- 02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md - 02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md
- 02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md - 02-DECISIONS/0126-a-module-declares-its-own-seats.md
- 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md - 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md
- 02-DECISIONS/0084-which-provider-serves-a-consumer.md - 02-DECISIONS/0084-which-provider-serves-a-consumer.md
- 02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md - 02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md
@@ -64,7 +64,7 @@ Which module answers, in order:
1. **the holder of the seat the requirement names.** A requirement may name a seat instead of leaving 1. **the holder of the seat the requirement names.** A requirement may name a seat instead of leaving
the provider open. It asks for *the mesh's* one, and the mesh answers with whichever assignment the provider open. It asks for *the mesh's* one, and the mesh answers with whichever assignment
holds that seat, with nothing asked of anyone. Unheld, the requirement is refused, naming the seat holds that seat, with nothing asked of anyone. Unheld, the requirement is refused, naming the seat
([ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md), ([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) (superseding [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md)),
[26 — The seats](26-the-seats.md)). Only a seat that delivers a provision can be named; naming a [26 — The seats](26-the-seats.md)). Only a seat that delivers a provision can be named; naming a
foundation seat is refused, because it delivers nothing. A `secret` requirement always names foundation seat is refused, because it delivers nothing. A `secret` requirement always names
`mesh-vault`, because that provision is reserved; `mesh-vault`, because that provision is reserved;
@@ -199,7 +199,7 @@ containers, login, broker account and settings are keyed by it, as today, and a
tightest backend ([ADR 0049](../../02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md)). tightest backend ([ADR 0049](../../02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md)).
**A module may run on many nodes, and one assignment may hold a seat** **A module may run on many nodes, and one assignment may hold a seat**
([ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md)). The definition ([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) (superseding [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md))). The definition
says which seats the module can hold; the assignment says which it does. So the store module can run says which seats the module can hold; the assignment says which it does. So the store module can run
on every node, one of those assignments holds `mesh-store`, and moving that role changes an on every node, one of those assignments holds `mesh-store`, and moving that role changes an
assignment, not a definition. assignment, not a definition.
+540 -45
View File
@@ -1,15 +1,21 @@
--- ---
layer: to-be layer: to-be
status: proposed status: in-progress
code: [] code:
updated: 2026-09-26 - mesh-catalog modules/nats
- mesh-controller internal/catalogue
- mesh-lab scenarios
updated: 2026-09-28
decisions: decisions:
- 02-DECISIONS/0116-the-bus-is-built-in-five-steps.md - 02-DECISIONS/0116-the-bus-is-built-in-five-steps.md
- 02-DECISIONS/0106-the-bus-is-nats.md - 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
- 02-DECISIONS/0126-a-module-declares-its-own-seats.md
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md - 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
- 02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md - 02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md
- 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md - 02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md
- 02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md - 02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md
- 02-DECISIONS/0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md
--- ---
# 28. Building the bus # 28. Building the bus
@@ -111,25 +117,153 @@ step 5 the rollout
## Step 1 — the module, and genesis raises it ## Step 1 — the module, and genesis raises it
> **Revised 2026-09-26** ([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md),
> [design 29](32-what-a-module-declares.md)). Tasks 1.3 and 1.4 said the controller composes every
> account and creates *the four streams* at genesis, from a fixed set. That is only the mesh's own
> half. A module declares seats with their protocols, so streams are created **at registration**
> and durable consumers **at assignment** — neither of which has happened at genesis. The fixed
> foundation set stays here; the derived machinery moves to step 3, where the declaration model it
> reads from is specified. Tasks 1.1 and 1.2, already done, are untouched by this: the module and
> its reload mechanism do not care what the configuration says.
**Why here.** Everything else needs a server to talk to, and genesis is where the foundation is **Why here.** Everything else needs a server to talk to, and genesis is where the foundation is
defined. The mesh this is for will never travel this path — it is already running, and takes step 2 defined. The mesh this is for will never travel this path — it is already running, and takes step 2
— but genesis is the definition every other path is measured against, and one that exists only on — but genesis is the definition every other path is measured against, and one that exists only on
paper is wrong until there is a second mesh to find out. paper is wrong until there is a second mesh to find out.
- [ ] 1.1 the `nats` module: manifest, image, one container, its client, TLS and monitoring ports, - [x] 1.1 the `nats` module: manifest, image, one container, its client, TLS and monitoring ports,
JetStream on a named volume — the shape of design 25 §5, and the same shape the broker module JetStream on a named volume — the shape of design 25 §5, and the same shape the broker module
beside it already has beside it already has
- [ ] 1.2 the composed configuration as a **directory** resource, and the entrypoint that watches - [x] 1.2 the composed configuration as a **directory** resource, and the entrypoint that watches
the one file and signals the server itself — design 25 §5's correction, kept inside the module the one file and signals the server itself — design 25 §5's correction, kept inside the module
because a container has no reload and a recreate would drop every connection the mesh has because a container has no reload and a recreate would drop every connection the mesh has
- [ ] 1.3 the controller composes that file: accounts, permissions, TLS, JetStream — permissions - [x] 1.3 the controller composes that file: accounts, permissions, TLS, JetStream — a user's
derived from `emits` and `consumes` and nothing else, plus each user's own ack subject and its permissions derived from its declaration and nothing else, over the three namespaces of
own inbox prefix (design 25 §4) [design 29](32-what-a-module-declares.md) §2, plus its own ack subject and its own inbox
- [ ] 1.4 the four streams, created at genesis and asserted idempotently on start, by the controller prefix (design 25 §4)
as their only writer - [x] 1.4 the mesh's own streams, created at genesis and asserted idempotently on start, by the
- [ ] 1.5 genesis raises it as foundation, claiming the seat **`mesh-broker`** — the seat is the controller as their only writer — **the mesh's own, not all of them**: a seat's streams are
server's role, not the product created when the module declaring it is registered, and a module's durable consumers when it
- [ ] 1.6 the genesis-broker bed is assigned, so this task is the fixed foundation set and 3.x carries the derived rest
- [x] 1.5 genesis raises it as foundation, claiming the seat **`mesh-broker`** — the seat is the
server's role, not the product. **Already true of the controller and needed no change**: it
resolves the broker by seat ("that is where the broker is, whatever else the topology says")
and names no broker module anywhere in its source. What remains is naming `nats` instead of
the deprecated broker where a genesis module set is declared, which is scenario and installer
configuration — carried with 1.6 rather than before it.
- [x] 1.7 **the composition, delivered** — the controller gathering its principals, composing the
file, and asserting the streams and consumers on start.
**In**: the user list is derived from the mesh's records and the credentials are kept.
A bus user's bcrypt hash is now recorded, keyed by the username the file needs, and the
plaintext is returned exactly once. That state is new and the reason is worth stating: on the
bus the mesh runs on today an account is a management call — mint, hand over, seal to the
holder, keep nothing — and that works because the broker remembers. Here the users are one
file rewritten whenever any of it changes, so keeping nothing would mean **the first person's
access change silently blanking every module's password**.
**Permissions are not kept, only credentials.** Authority is derived from what each module
declares every time the file is written ([ADR 0043](../../02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md));
a stored permission list would be a second account of a user's authority, able to disagree
with the records it came from while both looked internally consistent.
The derivation refuses two things where they can still be named: two users with one name — the
server reads the file as one of them and which one depends on the order — and a module
assigned but absent from the catalogue, which would compose a user with no authority and fail
on its first publish with an authorisation error that says nothing about a missing manifest.
A user the mesh has minted no password for is *named* rather than dropped or written as a user
anybody is: an ordinary situation with an obvious remedy, and the caller decides whether a
partial file is worth writing. A seat's protocol is gathered across the whole catalogue, not
from one manifest, because a seat is declared by one module and held by another.
**Out, and what each needs.**
**Delivery is in, and it settled what a module declares.** The mesh writes the *accounts* and
the module owns its *server*. The alternative was a manifest field enumerating ports, TLS
paths and a store directory so the controller could write a whole configuration — wrong,
because those are properties of the container the module raises and the controller would have
to be kept in step with a Dockerfile it never sees. So a module declares its own configuration
as a file resource and `bus-users` names where the mesh's half goes beside it; **asking is not
enough to receive it**, because that file holds every user's password hash, so the claim on
`mesh-broker` is what authorises it.
Two things a running server changed. **An absolute include path is resolved relative to the
including file's directory** — `include /etc/nats/accounts.conf` from another directory makes
the server look for it *under* that directory and refuse to start — so both files share one.
And **`verify: true` was refusing every connection in the mesh**: it makes the server demand a
*client* certificate, and nothing in the mesh presents one — a host pins this server's exact
certificate and authenticates with the password the mesh minted ([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md),
design 25 §4). Every connection would have died at the TLS handshake before any password was
looked at, with an error that reads as a fault in the client. Removed; TLS is still required,
because the block is what requires it and `verify` only decides whether client certificates
are checked. **Design 25 §4 should say this**, and says nothing about it today.
Also collected here: task 1.2's payoff, end to end against the module's own image — the user
list rewritten, the module noticing and reloading the server itself with no signal from
outside, and the connection the mesh already had still working afterwards.
**Minting is in, on both halves.** A node at enrolment, a module when its credential is
issued. Three things differ from a management call and each is the point of the move: the
credential is minted into the mesh's records and becomes usable at the next composition, so no
server need be reachable for it; the password travels beside the address rather than inside it,
because a credential embedded in a URL leaks into every log line that prints a connection; and
a module's durable consumer is derived from what it declared rather than named, so it cannot ask
for delivery of something it did not say it consumes. A node reconnecting may be refused until
the composition reaches the machine running the bus — which is what the host's reconnect backoff
is for, where waiting for the push would hold an enrolment open for as long as a declaration
takes to apply.
**Which bus is one fact, and being told about both is refused at start.** Not warned about: a
mesh half on each is one where a declaration goes out on one bus and the report comes back on
the other, and every component logs success while it happens — ADR 0074's failure arriving
through configuration instead of through code. A node that came away holding a credential for
each could be half-moved, and nothing would say which half.
**The objects are asserted on every start**, not created once at genesis: a stream somebody
deleted, a mesh raised from a restored backup, or a bus whose data directory was replaced all
have records and no objects, and a node whose consumer is missing hears nothing while everything
else about it looks correct. Against a real server: every object accepted, asserting twice
changes nothing (a start that failed the second time is a controller that cannot restart), a
machine joining an already-raised bus accepted, each node's consumer bound to its own
declaration subject and no other's, and CONTROL not dead-lettering — because the store window's
bound is the controller's, and a server that gave up first would discard the push the stream
exists to protect.
**People are not in the list**, deliberately: the account model is built and `operator issue`
is not (4.4), so there is nobody to derive. Left empty rather than guessed at.
> **This corrects a tick, not a decision.** Tasks 1.3 and 1.4 are ticked and they are honest
> about what they built — the composer, the derivation, the permission model, the stream and
> consumer definitions, the asserter, all pure and held by unit tests and a golden
> composition. What nobody wrote is the *caller*. Measured on the feature branch: outside the
> package that defines them, there is **not one** use of the composer, the permission
> derivation, the stream set, the stream asserter or the principal type. Step 1's "done when"
> claims "every account and permission composed from the manifests", and a mesh raised today
> would stand up a server with no user list at all.
>
> It also needs state the mesh does not keep. Design 25 §4 says the file holds bcrypt
> hashes, and passwords are "minted and sealed exactly as today" — but today the mesh mints a
> password, hands it to the broker through a management call, seals the plaintext to the
> holder and **keeps nothing**. There is no management call here, so the hash has to survive
> for every later recomposition: the first thing a person's access change or a new module
> touches is a file that must still contain every other user's password. No bcrypt hash is
> stored anywhere in the controller today.
>
> Named as its own task rather than folded into 1.3 so the gap is visible: the parts of
> step 1 exist and the mesh does not yet do any of it.
- [ ] 1.6 the genesis-broker bed — **deferred**: beds are run once, at the end, rather than per
step (novox/hq design 22's rule, and the operator's instruction). Every claim step 1 makes
is covered by a unit test or was demonstrated against the real server; what the bed adds is
the claims that need a mesh.
> **Not done here, deliberately.** The controller builds a module's broker credential as an
> `amqps://` URL and defaults a portless genesis address to 5671. Those are correct until the
> rollout and must not move: steps 1 to 4 leave every node on AMQP
> ([ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md)), so changing the
> credential's shape now would break the running bus to serve a bus nothing speaks yet. They
> change with the links, in step 3.
**Done when.** A mesh raised from nothing has the server standing with the streams asserted and **Done when.** A mesh raised from nothing has the server standing with the streams asserted and
every account and permission composed from the manifests; a user cannot publish outside its every account and permission composed from the manifests; a user cannot publish outside its
@@ -149,11 +283,35 @@ and it is what makes steps 3 and 4 safe to develop against a live mesh. A runnin
a foundation module by being raised again; it adopts one in place a foundation module by being raised again; it adopts one in place
([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)). ([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)).
- [ ] 2.1 the server raised beside the existing broker on its own ports, carrying nothing - [ ] 2.1 the server raised beside the existing broker on its own ports, carrying nothing —
- [ ] 2.2 the `nats` module adopted onto it in place, holding the data and configuration it was installer-side, from the upstream image
raised with - [ ] 2.2 the `nats` module assigned, which **recreates the container once, deliberately** (see
- [ ] 2.3 the seat claim, and the resolver's refusal of a second holder mesh-wide below), keeping its JetStream directory
- [ ] 2.4 the adoption bed - [x] 2.3 the seat claim, and the resolver's refusal of a second holder mesh-wide — **already
true and now proved**: the refusal is generic to any mesh-scoped seat, and three tests pin
what matters for this one — a second bus anywhere is refused naming the seat, a *different*
bus implementation is refused for the same reason (which is what lets the bus be replaced
at all), and the deprecated broker no longer contends for it, so both run on one mesh
- [ ] 2.4 the adoption bed — deferred with the other beds
> **Adoption here is not a no-op, and pretending it would be is the trap.** The host keeps an
> existing container only when its spec matches the declaration exactly
> ([`apply.go`](https://git.novox.be/novox/mesh-host): *existed && before.Spec == want && running*
> → unchanged; anything else is `rm -f` and recreate). Genesis raises the server from the
> **upstream** image, because nothing has been built yet; the module declares the **mesh-built**
> artifact, which carries the entrypoint that reloads configuration in place. Those two specs
> differ, so assigning the module recreates the container.
>
> That is correct, and it is [ADR 0067](../../02-DECISIONS/0067-genesis-is-a-pivot.md)'s pivot
> exactly: raise a temporary thing, then reinstall it as an ordinary module. It is safe **only
> because it happens while the bus carries nothing** — which is what 2.1 means by "carrying
> nothing", and why step 2 comes before anything speaks NATS rather than after. One recreate, at
> the one moment it costs nothing.
>
> **After that, never again.** The configuration is a directory mount rather than a file, so
> rewriting accounts does not change the container's spec and the entrypoint reloads the server in
> place. That is the whole point of task 1.2, and this is the moment it pays: every later account,
> permission or person's access change touches a running bus with connections on it.
**Done when.** A mesh already running has the server adopted, holding `mesh-broker`; a second **Done when.** A mesh already running has the server adopted, holding `mesh-broker`; a second
assignment anywhere is refused at resolution — *one per mesh*; and every node is still on the old assignment anywhere is refused at resolution — *one per mesh*; and every node is still on the old
@@ -166,19 +324,160 @@ quietly carried traffic would be step 5 arriving early and unrehearsed.
that decides what "agreeing" means for everything after it. It is the largest step and the one that that decides what "agreeing" means for everything after it. It is the largest step and the one that
pays for itself furthest away. pays for itself furthest away.
- [ ] 3.1 **the suite first, on the bus the mesh has** — fixtures for the envelope and its required - [x] 3.1/3.3 **the fixtures** — one directory in the sdk, read by each implementation's own
headers, the contributions file, a served tool call and a grant, capturing what the three runner rather than copied into either, because a fixture copied twice is two fixtures. The
implementations do *today*. Written first because a suite born on the new bus certifies Go emitter and the runtime's NATS client both pass the first: every required header set,
whatever the new bus happens to do each value in the pinned shape, the subject derived the same way, and the payload the body
- [ ] 3.2 design 19 rewritten from exchanges, queues and routing keys to the subjects and streams of alone.
**The suite also had to settle what "byte-for-byte" can mean**, which ADR 0074 stated and
nothing had yet had to implement. The envelope is exact — subject, required headers, names
and formats — because that is what two implementations get wrong invisibly. The body is
not: Go sorts a map's keys and JavaScript keeps insertion order, so identical bytes would
commit every implementation to a canonical JSON encoder, to buy a property the mesh never
uses. Read strictly it would have sent somebody writing one.
Still to capture: a served tool call, a grant and its answer, and the contributions file —
the other three ADR 0074 names.
- [x] 3.2 design 19 rewritten from exchanges, queues and routing keys to the subjects and streams of
design 25 §2–§3, per capability, with ADR 0074's model untouched: floor plus capabilities, an design 25 §2–§3, per capability, with ADR 0074's model untouched: floor plus capabilities, an
implementation legitimate when it claims less, identity from the sealed credential, dedup on implementation legitimate when it claims less, identity from the sealed credential, dedup on
`x-event-id` `x-event-id`. Claims checked against a running server are marked *verified* in the text, so
- [ ] 3.3 the fixtures restated on NATS, and the capability each implementation claims a reader can tell what was measured from what was reasoned. One limitation lifts with the
- [ ] 3.4 the controller's link on NATS transport: a module may now call another's tool, which issue 049 recorded it could not.
- [ ] 3.5 the host's link on NATS — mirroring, still importing nothing - [x] 3.4 the controller's link on NATS — **both halves are through the seam, and the store
- [ ] 3.6 the tool runtime's client on NATS, behind the unchanged sdk contract window is the server's.** `Bus` states the outbound in the mesh's words (publish an event,
- [ ] 3.7 the sdk's three stale comments, and nothing else in it declare to a node) and `Control` states the inbound (took it, dropped it, held it for the
store); each has an AMQP and a NATS implementation, and both ship, because steps 1 to 4
leave every node on AMQP and both shipping is what holds them to one envelope.
The outbound seam turned out to be eight call sites; the inbound was the larger half, and
the reason: every handler took the transport's own delivery type, so the loop could not move
without moving enrolment, reports, builds, upgrades and catch-up with it in one breath.
**The window (ADR 0083) is now what decides, once, for both.** On the bus the mesh has,
holding a message means an unacknowledged delivery kept in the controller, bounded by the
prefetch and lost if it stops. On the bus being built it is a `nak` with a delay: the
message stays the server's and the controller keeps only the moment it first could not take
it, so one that restarts mid-window has nothing to lose. Seven claims about that were asked
of a running server rather than reasoned — a report heard and gone from the work queue, one
held through a store outage and recorded when it returned, one let go once the bound passed,
a superseded one settled without being acted on, a heartbeat heard and nothing persisted,
both followed events acknowledged on a stream the controller had no ack subject for, and the
enrolment answer arriving at the address the request carried in its payload.
**Three things the wiring forced into the open.**
*Supersession is asked before the store, not after.* A report about a declaration the mesh
has moved past would otherwise wait out a restarting store to be written, and then overwrite
what the node is doing now.
*Half of a report is not about a declaration, and that half is never stale.* What the machine
**is** — the tunnel it took over, the ports its own bundle holds, what an adopted node found,
a node moving its overlay key — reaches the mesh on a report and nowhere else. A rekey set
aside as stale is a node whose overlay key never moves, and no retry is coming, because the
node said it once. So staleness is asked only of a report that is purely an apply's account.
*The controller could not have consumed a module event at all.* Its account granted no event
subject to subscribe and no ack subject on the events stream, so every announcement would
have been redelivered for ever, refused by the permission list it already had. Both are now
granted, each subject named rather than by pattern — a controller subscribing every event in
the mesh is a permission list that has stopped saying what it is for. Its consumers are
**named beside the mesh's own streams rather than derived**, because the controller files no
manifest and authority cannot come from a declaration that does not exist.
Still outstanding: a build's own shape, which travels with the builder in step 4.
- [x] 3.5 the host's link on NATS — **all three halves are through seams**, mirroring the
controller's and still importing nothing of the mesh's own (ADR 0005): the host's own
interfaces over its own libraries, agreeing with the controller only because a fixture holds
both to one envelope. A report goes through JetStream because it is the message the
store-window guarantee is about; a heartbeat stays on core, because a heartbeat in a stream is
the mesh's least valuable message competing for retention with its most valuable.
`Link` is dialling, hearing and saying in one interface, because **dialling is where the
transport is chosen** and choosing it twice is how one half of a node ends up on a different
bus from the other. `Asking` is the enrolment conversation, and it is separate for the
opposite reason: almost nothing about it is the same, and a node that fails there is not in
the mesh at all.
**What the new bus took away, and what it would not give.** A host declares nothing here: on
the old bus it declares its own queue, because a queue that is not there means a node that
hears nothing, but the object it reads through now is a durable consumer and a host's account
reaches no part of the JetStream API. So it **binds** to one the mesh made, and a missing one
is said as the mesh's to answer rather than quietly created with whatever the client defaults
to. Two things that had to be built for that: a node's declaration consumer (named after the
node, because its ack grant is derived from the node's name, so any other name is a delivery
it cannot acknowledge), and the enrolment user's **inbox** — design 25 §6 names it and the
composer granted none, so an enrolling node would have published its request and waited out
its timeout against a mesh that answered.
**The reply address travels in the payload, and that is now proved from both ends.** The
controller reads it from there (3.4) and the host writes it there and waits on it, and the
test asserts the transport's own reply field held the *consumer's ack address* by the time the
request arrived — so a future server that stopped claiming that field fails a test rather than
letting the reason quietly become folklore.
The host's **"newest wins" window narrows at the rollout rather than disappearing**, and that
is now measured rather than predicted: three declarations pushed to an absent node leave one
on the stream and it is the newest, so the catch-up half is the stream's — but three pushes to
a connected node are still three deliveries, which is the half that stays.
**The pin turned out easier here than in the tool runtime, not harder.** The Go client takes a
`*tls.Config`, so the same pinned configuration with the same verify callback does the work;
the subject-alternative-name constraint recorded under 3.6 is that client's, because it takes
PEM strings with no verify hook. A host checks the fingerprint and nothing else.
Nothing here composes an enrolment user per live token, and that is **1.7's**, not this
task's: it is one input to a composition that does not happen at all yet.
- [x] 3.6 the tool runtime's client on NATS, behind the unchanged sdk contract — round-tripped
against a real server: a tool answered across two connections, a throwing handler reaching
the caller as an error rather than a timeout, an event delivered once with its key, body,
node and event id intact. Ships beside the AMQP client and is selected at the rollout,
because steps 1 to 4 leave every node on AMQP.
**A constraint it surfaced, recorded where somebody issuing a certificate will look.** The
AMQP client pinned the exact certificate and switched hostname verification off, which is
sound because a fingerprint is stronger than a name. The NATS client exposes no equivalent
hook — its TLS options are PEM strings with no verify callback — so the pin still happens
before dialling and the library's own name check happens beside it. **The bus's certificate
must carry a subject-alternative name matching the address nodes dial it by**, or the
connection is refused by a library error rather than by anything the mesh says.
- [x] 3.7 the sdk's three stale comments, and nothing else in it — three lines, which is the
whole of the sdk's diff for the bus change, and the measurement that predicted it
- [x] 3.8 **the declaration model** of [design 29](32-what-a-module-declares.md): local names
derived to subjects, the three namespaces, permissions computed from a declaration, and a
manifest that contains no subject. Done in the controller's composer (permissions, streams,
consumers), in the runtime's client (subjects derived from the credential, never named by a
module), and as a catalogue test asserting all 72 manifests hold no subject — because the
rule held by construction, and a rule held by construction is one a later field breaks
quietly.
- [x] 3.9 **seats declared by modules** — the manifest now carries `seats` (name, scope,
accepts/emits/serves, retention) and `uses`, and registration refuses a `mesh-*` name, a
duplicate declarer, an undeclared `uses` or claim, a seat with no protocol, a scope
mismatch, and a holder that does not answer what its seat promises. **Still to do**:
creating a seat's streams at registration and its holder's work-queue consumer at
assignment, which need the JetStream client wired in.
The refusal for an unknown claim *moved* rather than disappeared — the parser cannot judge
it from one manifest any more, because another module may legitimately declare that seat,
so it is registration's. The test that encoded the old rule was rewritten rather than
deleted, and a second one pins the case the parser could not distinguish.
**Done**: a seat's work queue is derived and created, and a holder's worker with it. The
JetStream client behind them is wired and verified against a running server, which also
completes 1.4's missing half — the pure `Asserter` had no implementation until now.
- [x] 3.10 **the ten seat renames** — done in the controller's table, the ten manifests that
claim them, the controller's own shipped manifests, and every test. Not a migration after
all: a holding is derived at resolution, never stored, so nothing recorded points at an old
name (recorded as a progressive insight on ADR 0126). A **kept** rename table tells a
manifest written against an old name what it became, because a module lives in its own
repository and may be registered long after the catalogue stopped using one.
**A seat and the interface it delivers are different names.** The `git` seat became
`mesh-git` while the `git` *provision* it delivers did not change, and the same for the
package registry. A blanket replace got this wrong first and the failure read "the package
registry is served on `<nil>`", which does not say "you renamed an interface" — so a test
now pins every seat against the interface it delivers.
**Done when.** The fixtures are produced and consumed byte for byte by every implementation that **Done when.** The fixtures are produced and consumed byte for byte by every implementation that
claims the capability, and a module built before any of this serves its tools unchanged on the new claims the capability, and a module built before any of this serves its tools unchanged on the new
@@ -196,6 +495,17 @@ module.
**Why here.** The links exist from step 3, so the flows that are not on the bus at all can move onto **Why here.** The links exist from step 3, so the flows that are not on the bus at all can move onto
it, and the beds that need a mesh living on NATS can finally run. it, and the beds that need a mesh living on NATS can finally run.
> **A blocker surfaced here that is not this step's to fix.** Every event name in the catalogue is
> still written the way a routing key on the bus the mesh has is written, so the derivation design 29
> §1 specifies turns a consumer's declaration into a subject **no emitter publishes** — thirty-seven
> manifests, and one that cannot be composed at all. Nothing fails on the bus the mesh runs on
> today, where a routing key is matched literally; it fails on the first mesh raised on the new bus
> and not before, which is why wiring the controller's own subscription is what found it. Opened as
> [issue 127](../../04-ISSUES/127-a-module-event-derives-a-subject-nothing-publishes/00-report.md).
> It holds 4.2, 4.3 and the catch-up half of 4.5; the node-facing flows — enrolment, reports,
> heartbeats, a build's outcome — are unaffected, because those subjects are the mesh's own and
> derive from nothing a module declares.
- [ ] 4.1 **the full genesis bed** — a mesh raised on NATS from nothing and living on it: a node - [ ] 4.1 **the full genesis bed** — a mesh raised on NATS from nothing and living on it: a node
enrols over TLS with a claimed token and the enrolment user cannot read a declaration; a push enrols over TLS with a claimed token and the enrolment user cannot read a declaration; a push
is held while the store restarts and applies after, nothing lost or duplicated; a node that is held while the store restarts and applies after, nothing lost or duplicated; a node that
@@ -204,13 +514,111 @@ it, and the beds that need a mesh living on NATS can finally run.
held by a `nak`-with-delay cycle still reaches the enrolling node, proving the reply travels held by a `nak`-with-delay cycle still reaches the enrolling node, proving the reply travels
in the payload and not the transport field the consumer's ack has claimed. The server-enforced in the payload and not the transport field the consumer's ack has claimed. The server-enforced
permissions were proved at step 1 and are not re-proved here permissions were proved at step 1 and are not re-proved here
- [ ] 4.2 a build source's change reaches the builder over the bus, and the build that follows is — **nothing is outstanding but the bed itself.** Both links speak NATS, the composition
the one the change asked for happens, and every claim above has a unit test or a check against a running server behind it.
- [ ] 4.3 an installation completes over the bus, with the same outcome as the path it replaces What none of them can stand in for is a mesh raising itself, which is what this bed is — so this
- [ ] 4.4 a person's client: the account, the client that speaks the bus, and the tool surface over is where the code stops and the lab starts
it (design 25 §7) — a module's tool invoked from another node and from a person, refused from - [x] 4.2 a build source's change reaches the builder over the bus, and the build that follows is
an account that may not the one the change asked for — **a build is work submitted to a role now**
- [ ] 4.5 reports and catch-up: a node that was unreachable catches up rather than losing them ([ADR 0129](../../02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md)). Both sides
are behind a seam with an implementation per bus, and on the bus being built one publish does
what two did: the outcome is the role's own event, so the asker matches it by the id its
request carried, the controller records it and the catalogue places it in the graph. A build
machine needs a reply queue for nothing and a grant over nobody's inbox.
Checked against a running server: the round trip; a third party on the role's event hearing
the same outcome the asker did, which is what the decision rests on; work leaving the queue
once settled, so no second machine repeats it; work submitted with no machine holding the role
**waiting rather than failing**, and being done when one arrives; and work a machine handed
back coming round again.
The outcome carries the module name, because only the manifest says what was built and one
message now has three readers. A failed build names none: it produced no module version, and
the catalogue would otherwise place something that was never made.
- [x] 4.3 an installation completes over the bus, with the same outcome as the path it replaces —
**the installer can raise it**: a foundation template that stands up the server, writes the
server's own settings and the mesh's first user list beside them, and starts a controller
reaching the new bus. What remains is running it, which is 4.1's bed.
**The mesh composes its own user list, and at genesis there is no mesh to compose one.** So the
installer carries the first — the controller's account at a well-known bootstrap password,
exactly as the store is reached at `postgres:bootstrap` and the old bus at `guest:guest`, and
rotated with them. From the controller's first composition onward the file is the controller's.
That surfaced a gap reading would not have found: the controller's own account exists before
there is a controller to mint one, so nothing recorded a hash for it and its first composition
would have left the writer out of the file it was writing — a bus nothing can connect to,
produced by the thing connected to it. It records a hash of the credential it is using, and only
when none is recorded, so a restart cannot put the bootstrap password back over a rotated one.
**The carried list and the derived one are checked against each other**, because they are two
statements of one fact and a mesh cannot be raised twice to find out they disagreed. A template
granting less than the controller derives produces a mesh that comes up, connects, and is
refused on its first act, with an authorisation error naming a subject rather than the template
that forgot it. The check earned itself at once: the composer was granting a role's whole event
branch *and* the one event it follows, and the wider grant wins — so only the submitting half of
a role is granted now, and what comes back is named exactly.
- [x] 4.4 a person's client — **the account and the program are both in.**
**The account**: a person is not a module and holds no seat, so their authority is a list of
tools (or `*` for an administrator) and nothing else. Held to four properties, each a way of
being wrong that would not announce itself: nothing but tools, so a person cannot claim a
module said something; no ack subject, because authority over a consumer that does not exist
is authority nobody audits; no ability to answer, because a person who can answer a request is
impersonating a module on a bus where anyone may serve a tool; and two people do not share an
inbox. Issued, listed and revoked by command; stating what somebody may call replaces what was
there, because a list that could only grow is a permission nobody can take back; and forgetting
somebody takes their credential with them, or it is not a revocation.
**The program**: two surfaces over one thing — a command line and an MCP server — both adapters
over the same three calls, because a second way of reaching a tool is a second thing to keep
correct. It uses the client a module's runtime uses, so what a person may do is answered by the
same permission list that answers it for a module and an audit has nothing separate to read.
Three decisions in it worth keeping. It lists what the **catalogue** has rather than what this
credential may call: somebody seeing only their own tools cannot tell "not installed" from "not
yours", and those need different people to fix them. A failed call says which of three things
happened — nobody serves it, this credential may not, or the tool was slow — because the
remedies are in three different places and without that they are one timeout and a stack trace.
And the MCP surface decides nothing: the names are the ones a person types, the schemas are the
modules' own, an answer is passed through unshaped, and a tool that fails comes back as a tool
error rather than a protocol error, because the request was well-formed and the mesh answered it.
Both surfaces are driven against a running bus, including a host's notification being answered
with nothing and an unknown method refused.
> **Design 25 §7 says "nothing is built of this before §10's bed passes", and this was built
> before.** Recorded rather than quietly ignored: the operator asked for it, it is on the
> critical path for nothing and blocked by nothing, and the bed it waits for is 4.1's. If the
> bed changes what a person's client should be, this is what gets changed.
- [x] 4.5 reports and catch-up: a node that was unreachable catches up rather than losing them.
**The reports half is in and proved against a server** (3.4): held through the store's absence by
the server rather than by the controller, superseded ones settled by the digest they carry.
**The catch-up half needed nothing built, and that was the answer.** It existed because a queue on
the bus the mesh runs on today receives only what is published after it is bound, so everything
built before the catalogue existed was announced to nobody — and on a fresh mesh that is always
the foundation, because those are the things the catalogue needed in order to exist
([issue 050](../../04-ISSUES/050-the-catalogue-knows-nothing-built-before-it/00-report.md)). A
whole mechanism followed: the catalogue asks, the controller re-publishes.
A stream is a log and a consumer is a position in it. A consumer created later starts at the
beginning, so the builds are simply there — asked of a running server rather than assumed, since
the decision rested on it: three builds published with nothing listening, then a consumer created,
and all three waiting for it. So the question of *who replays* has no answer because nothing
replays.
> **This is the shape of the whole change, in one task.** Three ways to do the replay were weighed
> — a namespace for the mesh's own voice, the controller answering a question, a consumer reading
> from the start — and the right answer was that the bus being moved to already does it. The
> mechanism was never about builds; it was about a queue that could not remember. **A conversion
> that carried it across would have carried a workaround for a limitation that no longer exists**,
> and nothing would have looked wrong.
Retiring it is step 5's, with the rest of what only the old bus needs: the request, the
re-publishing, and the `replay` flag that told a consumer to register history without acting on it.
**Done when.** Each converted flow is proved against the behaviour it replaced, and the full genesis **Done when.** Each converted flow is proved against the behaviour it replaced, and the full genesis
bed is green. **Observation is not in this step** — heartbeats, conditions and key-value state are bed is green. **Observation is not in this step** — heartbeats, conditions and key-value state are
@@ -221,16 +629,101 @@ reserves them for after the move, and a flow built ahead of its design would be
**Why here.** It is the only step that moves a node's bus, and it moves every node's at once. **Why here.** It is the only step that moves a node's bus, and it moves every node's at once.
- [ ] 5.1 the cutover bed: a mesh on AMQP with a predecessor stand-in on the compatibility broker **The order the repositories land in is part of the rollout, not paperwork.** Derived 2026-09-27
while merging, and not obvious from any one repository, which is why it is written here rather than
left to be re-derived under time pressure:
| order | repository | why it cannot be later |
|---|---|---|
| 1 | this one | prose; nothing deploys |
| 2 | the sdk | comments only, and no module rebuilds for it |
| 3 | the client library | it is what a module calls to emit, and it is where the subject is derived. Until it lands, a locally-named event is published under the local name itself |
| 4 | the catalogue | every manifest and every module's code, renamed together. Safe only once the runtime derives |
| 5 | the controller | **it refuses an old-style event name outright**, so landing it before the catalogue makes every unconverted module unregisterable |
| 6 | the hosts | last, because nothing else waits on them |
Two properties make the sequence safe rather than merely ordered, and both are pinned by tests. A
name already in the old form passes through the derivation untouched, so a module nobody has
converted keeps working at every step. And a converted name derives to **exactly** the key the old
bus published, so steps 3 and 4 change nothing on the wire — the move to the new bus is step 5.2 and
one environment variable, not a side effect of deploying.
The failure this ordering avoids is issue 127's own: a publisher and a subscriber that disagree about
a subject produce no error anywhere. Nothing logs, nothing retries, and the mesh reports itself
healthy while reacting to nothing.
- [ ] 5.1 the cutover bed: a mesh on AMQP with a predecessor stand-in on the deprecated broker
moves its bus in one rollout, every node reporting on NATS afterwards, the stand-in's own moves its bus in one rollout, every node reporting on NATS afterwards, the stand-in's own
client still connected throughout client still connected throughout
- [ ] 5.2 the rollout: accounts composed, then the controller, every host and every runtime - [x] 5.2 the rollout: accounts composed, then the controller, every host and every runtime
together; every node confirmed heard before AMQP stops together; every node confirmed heard before AMQP stops.
- [ ] 5.3 the mesh's accounts removed from the compatibility broker, leaving the predecessor's users
- [ ] 5.4 the compatibility broker retires when its condition holds — no client connected for the
period the operator sets
**Done when.** Every node reports on NATS and the predecessor's clients never noticed. **Done 2026-09-28, 02:25.** Every machine reports on the new bus, the seat is held by the
module that provides it, the old broker is unassigned and forgotten, and every credential was
minted afresh at the end because two had been printed on the way. What it took, in the order
it was found, each fixed on the trunk before the next step: the control plane's `serve` and
`push` never selected the new transport (task 4.3, open until then); a machine's user was
granted neither the asking nor the delivery of its own consumer; the account had no JetStream
of its own; the control plane's client verified the bus's certificate by name instead of
pinning it; the seat table's rows carried no protocol, so no role's work queue was raised; the
build machine decided its bus from a variable its container never received; and a rotation
put new hashes on the bus before three machines had received their new memberships — which
is why there is now `rollout hand <node>` and a host adopts a delivered membership at start.
The bootstrap loop — a bus that can only be raised by a declaration that can only arrive
over that bus — was broken once, by hand: the mesh's own composed configuration started the
server, and the controller binary was run on the node directly until the managed container
could be rebuilt over the bus it was on.
- [x] 5.3 **the seat changes hands as one act.** A command takes a seat and the assignment taking it
over, and the seat is never empty in between — the emptiness is the outage of 2026-09-27, when
the control plane, which finds its own bus through this seat, lost the address and looped.
**Built 2026-09-27** (`seat_holding`, migration 0039; design 26 says how it is checked), and used
live the next night to hand `mesh-broker` from the old broker's assignment to the new one's. This
is what 5.2 uses to move `mesh-broker` from the old
broker's assignment to the new one's, and it is built first ([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)).
- [x] 5.4 **the old broker and everything that named AMQP leave the mesh** ([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md),
superseding [ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md)): the two modules that
required `amqp` are removed, the broker's module is unassigned and removed (**done 2026-09-28**; the predecessor's own tooling, which rode the same adopted broker, went dark with it, as [ADR 0130](../../02-DECISIONS/0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md) accepted), registration refuses
a manifest that provides or requires `amqp`, and a whole-catalogue check asserts none does. Not
a retirement condition — a decision, taken, with the operator's "I don't care if the predecessor
breaks" on record ([ADR 0130](../../02-DECISIONS/0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md)).
Retiring with it: the build outcome's second announcement under the module's own name, which
existed only so a catalogue deployed before the rename and one after both heard it.
> **The remote tooling goes with it too.** The predecessor's own mesh talks over that broker, so
> shutting it down ends the path that reaches this installation's machines from a workstation.
> The rollout is driven from the node, or before the broker stops — a sequencing constraint on
> 5.2, not an afterthought.
- [x] 5.5 **the AMQP transport is deleted from the control plane and the hosts**. One bus, nothing
to select ([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)).
**Done 2026-09-28.** The control plane's old consume loop, build request, tool call, management
API and account scoping went, and the host's old dialling and enrolment paths with them; a
membership or token naming any other bus is refused before anything is sent. Nothing selects a
transport any more: the variable that once did (`MESH_BUS_NATS`) now only names where the
control plane reads its own bus credential, the way any module reads a secret. **Checked by the
build**: neither repository's module file names the AMQP client library, so a line that still
used it would not compile. The store-window guarantee ([issue 083](../../04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/00-report.md))
is tested against a bus-less fake rather than the old transport's memory, which is what let
that memory go — the one thing it did that the stream does not (superseding a held report) is
the staleness check on the message itself (design 25 §3).
Found on the way: **no build had ever recorded what it stood on.** A recipe reads its base from
a build argument, so the digest was never in the file the builder derived edges from, and every
order that says *bases first* — `build --on`, `build --behind`, the merge follow-up of
[issue 131](../../04-ISSUES/131-nothing-tells-the-mesh-a-source-moved/00-report.md) — walked a
graph with no edges. The builder now reports the bases it was handed, the control plane records
them by artifact path, and the graph is read from the newest build of each module — a recorded
manifest carries no `build.on`, so the edge is derived from the build or it does not exist.
> **The old 5.4 note is history.** It recorded that a retirement *condition* was wrong from
> [ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md) onward, which framed the old broker
> as an ordinary provider with no end. [ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md) ends that
> framing in turn: the broker is not kept as a provider either, because AMQP is not a provision. Both
> readings are kept here so the two reversals can be read in order.
**Done when.** Every node reports on NATS, and nothing of the mesh's own is left connected to the
deprecated broker.
## The through-line ## The through-line
@@ -245,8 +738,10 @@ itself moves once, at the end, on one day.
- **Observation** — research 017's, after the move, by its own design. - **Observation** — research 017's, after the move, by its own design.
- **Leaf nodes** — design 25 §11 keeps this out of scope and says so; a leaf per machine is a later - **Leaf nodes** — design 25 §11 keeps this out of scope and says so; a leaf per machine is a later
question, noted so it is not forgotten. question, noted so it is not forgotten.
- **The predecessor's world.** It is AMQP, it cannot move, and it does not need to: its broker is - **The predecessor's world.** It is AMQP and it is not moving —
the compatibility module until its last client is gone. [ADR 0130](../../02-DECISIONS/0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md): it is
deprecated, some of it is still running, and it is being left to stop rather than migrated. Its
broker goes with it, unassigned like any provider whose provision nothing requires.
## How this list is kept true ## How this list is kept true
@@ -5,7 +5,10 @@ code: []
updated: 2026-09-27 updated: 2026-09-27
decisions: decisions:
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md - 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
- 02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md
- 02-DECISIONS/0051-shared-data-is-the-operators.md - 02-DECISIONS/0051-shared-data-is-the-operators.md
- 02-DECISIONS/0113-the-vault-makes-every-secret.md
- 02-DECISIONS/0117-a-machines-uplink-is-a-seat.md
--- ---
# 29 — A node has operator accounts, and the mesh owns what lives under a home # 29 — A node has operator accounts, and the mesh owns what lives under a home
@@ -18,7 +21,7 @@ owned by, who a user service runs as, and — the case that surfaced this — wh
person's `~/.ssh/config`, `~/.zshrc`, `~/.config`); the mesh, taking those over, kept the machine person's `~/.ssh/config`, `~/.zshrc`, `~/.config`); the mesh, taking those over, kept the machine
facts and dropped the human one. facts and dropped the human one.
Two things are missing, and they are one idea: Several things are missing, and they are one idea.
## 1. The account is a node fact ## 1. The account is a node fact
@@ -28,17 +31,14 @@ mesh already knows the node and its address, so `<account>@<node>` is then a com
because it is exactly the fact that was silently lost — `ssh ace` failed to `ace` because nothing because it is exactly the fact that was silently lost — `ssh ace` failed to `ace` because nothing
in the mesh said ace's account is `ace`. in the mesh said ace's account is `ace`.
It is **not** a credential. The account names a login; the key that authorises it is the
operator's, placed as a secret or an operator-owned file, never minted by the mesh (ADR 0051).
## 2. A resource may live under a home, owned by its account ## 2. A resource may live under a home, owned by its account
[ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) placed a [ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) placed a
module's *system* data — `<root>/<module>`, owned by the module. It has no analog for the other module's *system* data — `<root>/<module>`, owned by the module. It has no analog for the other
half of the filesystem: the things that belong under a person's home and are owned by that half of the filesystem: the things that belong under a person's home and are owned by that
person. `~/.ssh/config`, `~/.ssh/config.d/mesh`, `~/.zshrc`, `~/.config/hal` — every one of these person. `~/.ssh/config`, `~/.zshrc`, `~/.config/hal` — every one of these is a resource the mesh
is a resource the mesh should be able to place and own, resolved against **the account's home** should be able to place and own, resolved against **the account's home** rather than a system
rather than a system root, and chowned to **the account** rather than to root or a module uid. root, and chowned to **the account** rather than to root or a module uid.
This is the same move as `${dir:…}`, one level over: a resource says `home: <account>` (or names This is the same move as `${dir:…}`, one level over: a resource says `home: <account>` (or names
an account requirement), and the mesh resolves the home directory and the owning uid on the node an account requirement), and the mesh resolves the home directory and the owning uid on the node
@@ -46,39 +46,133 @@ that account lives on. A module that writes operator config — the eventual rep
`hal/terminal`, `hal/claude-code`, `hal/secrets` — declares its files this way and names no `hal/terminal`, `hal/claude-code`, `hal/secrets` — declares its files this way and names no
`/home/...` path, exactly as a system module now names no `/var/lib` path. `/home/...` path, exactly as a system module now names no `/var/lib` path.
These are a **family**, not one module: an `ssh-client` module, a shell module, a `~/.config`
module, each a *universal-tier* consumer of the account fact — assigned wherever a person logs in,
which is every node, unlike the graphical stack that a capability gates.
## 3. The whole of `~/.ssh` is the mesh's — with one boundary drawn inside it
The predecessor owned a single file (`~/.ssh/config`) and left the rest alone; it drifted, because
owning one file beside foreign ones is not owning anything. The mesh should own **the directory**:
create `~/.ssh` at `0700`, chown it to the account, and own the files it places there —
- **`config`** (or the mesh's region of it): the `Host` blocks for every other node, composed
from the roster;
- **`known_hosts`**: authoritative, so the "Host key verification failed / accept-new" dance that
cost real time during enrolment simply ends;
- **`authorized_keys`**: who may log into this account, governed centrally rather than by whichever
key happened to be pasted where.
**The boundary — and it is the reason this is safe:** `~/.ssh` is the one directory where a wrong
declaration locks a person out of their own machine. So the mesh's *found-vs-owned* semantics
([ADR 0126](../../02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md),
adoption) apply *inside* the home directory. The mesh **owns** the directory and the files above; it
**holds as found — never rewrites, never removes** — the operator's own contents: their **private
keys** and their **personal drop-ins** (`config.d/personal`, the personal `Host` aliases a
workstation carries, exactly as `hosts.local` is the home the mesh never rewrites for `/etc/hosts`).
Reconcile removing an unassigned `config.d/mesh` is fine; the same logic aimed at `id_ed25519` or an
operator's own `authorized_keys` entry is a lockout. This is the login-channel cousin of the rule
[ADR 0125](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md) draws for the uplink and the sshd
module draws for the firewall: **the mesh must never be able to arrange the one failure that severs
its own way back in.** The carve-out is not a convenience; it is that rule, in `~/.ssh`.
## 4. Keys are the mesh's to generate — through a CA, and existing keys are adopted, not replaced
Key *generation* is the mesh's, not each node's improvising its own. The clean form is an **SSH
certificate authority as a seat**, the sibling of the TLS internal CA the mesh already runs:
- **Host certs.** The mesh signs each node's host key. Every node's `known_hosts` becomes one line
— `@cert-authority *.<suffix> <mesh-CA-key>` — and nothing is distributed per node; a new node is
trusted the instant its host key is signed.
- **User certs.** The mesh signs a cert naming the principals (accounts) allowed. Every node's
`authorized_keys` / sshd `TrustedUserCAKeys` becomes one trust line — no N×N key spraying — and
short-lived certs give rotation for free
([ADR 0114](../../02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md)).
- The **CA private key is the mesh's**, a secret the vault makes
([ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md)).
**Three kinds of key, and only one is never minted.** Host keys (server identity) and pure
machine-to-machine keys the mesh may generate end to end. The operator's **personal** private key —
possibly on a hardware token, possibly used from an off-mesh laptop — the mesh **signs into a cert
but never generates**; that, and only that, is the residue of
[ADR 0051](../../02-DECISIONS/0051-shared-data-is-the-operators.md). So "keys are mesh-owned" and
"the operator's login key is the operator's" reconcile: the mesh owns the CA and the signing; it
holds the human's private half, never mints it.
**Existing keys are not lost.** Taking ownership is *adoption*, not regeneration: a key already on a
machine is recorded and signed, not overwritten. The mesh gains authority over `~/.ssh` — it does
not clear it. An enrolling node's host key and the operator's existing key are carried forward; the
found-vs-owned boundary of §3 is exactly what guarantees nothing already there is destroyed.
## 5. How it is distributed: the controller composes, the node applies
None of this needs a node to discover the mesh, and none of it needs a control-plane module of its
own. The ssh files are **roster facts**
([ADR 0128](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md)): once the
roster view carries a node's **host key** and its **account** beside its name and address, the
`ssh-client` module ships a template for `known_hosts`, `config` and `authorized_keys`, and the
controller renders each node's copy from the full roster and pushes it. The mesh owns the data; the
module owns ssh's format; the control plane gains no ssh syntax. It is the same act as composing a
peer list or `/etc/hosts` — which is why there is **no novox-only "mesh-ssh" module**: the
centralization is the controller's composition, not a module that runs somewhere. Only non-secret
facts travel (names, addresses, accounts, host keys, the CA public key); the private key stays the
operator's, placed as an operator-owned file, referenced by path.
## 6. The two modules, and the seat between them
- **`sshd`** (server, every node) — manages sshd, owns and **reports** its host key so the roster
carries it, and trusts the user CA.
- **`ssh-client`** (client, every node) — owns `~/.ssh` per §3, consumes the roster and the CA
public key.
- **`the-ssh-ca`** (a seat, held on the control node) — signs host and user certs.
They meet at the account and the CA, not at a bespoke module. The `sshd` server side already exists;
the client/identity side and the CA are the open pieces.
## Why now, and why not yet ## Why now, and why not yet
**Why it matters:** when HAL retires, the generators that keep `~/.ssh/config`, shell config and **Why it matters:** when HAL retires, the generators that keep `~/.ssh`, shell config and the
the operator's `~/.config/hal` current retire with it. Without this, adding a node stops adding operator's `~/.config/hal` current retire with it. Without this, adding a node stops adding its ssh
its ssh alias, and a fresh machine has no operator dotfiles at all — the mesh would run every alias and its trust, and a fresh machine has no operator dotfiles at all — the mesh would run every
service and leave the human unable to work on the box. The account is also load-bearing for service and leave the human unable to work on the box.
correctness already: `ssh <node>` (issue 122's cousin), user-scoped systemd units, and any file a
person rather than a daemon must own.
**Why not build it reflexively:** it is a real addition to the node model and the resource model, **Why not build it reflexively:** it is a real addition to the node model, the resource model, and
and it must be gotten right, not smuggled in beside a firewall fix. Open questions to settle the seat set, and must be gotten right. The mechanism half is now settled — ADR 0128 is what lets
first: the ssh files be templates with no control-plane format — so what remains to decide here is the
model:
- **One account or several per node?** A workstation has one human; a shared box might have more. - **One account or several per node?** A workstation has one human; a shared box might have more.
The model should allow more than one without forcing the common case to name it. Allow more than one without forcing the common case to name it.
- **Where the login key lives.** An operator-owned file (ADR 0051) or an accepted secret — never - **The CA's shape.** Host-cert and user-cert principals, cert lifetime and renewal, where the CA
minted. The account fact and the key that authorises it are separate, and only the first is the runs (a seat on the control node). The one thing fixed: the operator's personal key is signed,
mesh's to generate. never minted.
- **The boundary with `sshd`.** The `sshd` module (server side) already exists. This is the - **Adoption of existing keys.** How an enrolling node's host key and an operator's existing key are
*client* and *identity* side: the account a node offers, and the home-scoped files an operator recorded and signed rather than replaced — the found-vs-owned boundary, made concrete for keys.
needs. They meet at the account but are not the same module. - **The `sshd` boundary.** Server side exists; this is the client, the identity, and the CA.
- **Multi-operator.** Today there is one human. The model should not assume it, but the first - **The ssh-agent.** An agent is a *user-scoped service running as the account* — the first concrete
cut may serve one and leave the shape open. case of the user services §2 anticipates. It holds the operator's private key in memory; the mesh
declares the unit and sets `AddKeysToAgent`/`IdentityAgent` in `config`, and still never sees the
private half. Agent *forwarding* wants a policy, not a default: with user certs it is largely
unnecessary, and forwarding an agent into a node exposes the operator's keys to that node's root —
so prefer certificates and `ProxyJump` over forwarding.
**Not urgent, not blocking.** ssh and dotfiles work today because HAL's generators still run as **Not urgent, not blocking.** ssh and dotfiles work today because HAL's generators still run as the
the substrate. This becomes load-bearing in the node-by-node retirement phase, not before — which substrate. This becomes load-bearing in the node-by-node retirement phase, not before — which is the
is the right time to build it, once the account model is decided here. right time to build it, once the account and CA model are decided here.
## References ## References
- The gap was found generating `~/.ssh/config` from the *HAL* registry (`hal/terminal`'s - The gap was found generating `~/.ssh/config` from the *HAL* registry (`hal/terminal`'s
postConfigure hook), which the nox mesh has no equivalent for. postConfigure hook), which the nox mesh has no equivalent for.
- [ADR 0128](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md) — the roster
fact mechanism that renders the ssh files, format owned by the module.
- [ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) — the - [ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) — the
system-path placement this mirrors for home paths. system-path placement this mirrors for home paths.
- [ADR 0051](../../02-DECISIONS/0051-shared-data-is-the-operators.md) — why the login key stays - [ADR 0051](../../02-DECISIONS/0051-shared-data-is-the-operators.md) — why the operator's personal
the operator's, never minted. key is signed, never minted.
- [ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md) — the CA key is a secret the
vault makes; [ADR 0114](../../02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md)
— short-lived certs as rotation.
- [ADR 0125](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md),
[ADR 0126](../../02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md) —
the never-sever-the-channel rule and the found-vs-owned semantics, applied here to `~/.ssh`.
@@ -0,0 +1,137 @@
---
layer: to-be
status: proposed
code: []
updated: 2026-09-27
decisions:
- 02-DECISIONS/0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md
- 02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md
---
# 30 — The mesh updates itself on a push
**Today the mesh does not update itself; a person drives the pipeline by hand, and one class of
change freezes it.** A code change lands in `mesh-controller` or `mesh-catalog`, and getting it onto
the machines is a sequence somebody types. The predecessor's pipelines rebuilt and redeployed on a
push without anyone watching; the successor should too. This records the process as it is done by
hand now — so it can be read, and then coded — and the two things that make it more than "add a
webhook".
## The process, as done by hand
**An ordinary (non-breaking) change** — new module code, a bug fix, a manifest tweak that changes no
seat or schema:
1. `module moved <module> <commit>` — tell the mesh its source advanced (the controller repo has no
trigger, so this is manual; the catalogue's webhook does it automatically — see below).
2. `build --behind` (or `build <repo> [--ref] [--path <subdir>]`) — the build machine rebuilds and
records the new image.
3. The mesh **reconciles on its own**: the module's declaration now names the new image, the next
push/heartbeat sends it, and the host swaps the container. For the control plane this is a
self-upgrade — the running controller composes its own new image and the host replaces it. No
restart is typed.
**A breaking change** — a manifest schema the controller parses differently (a fact's shape, a
seat's name), where the new control plane cannot read the manifests the old one stored:
4. Land the code (controller + catalogue together — they are one change).
5. Rebuild + deploy the new controller (steps 1–3). **The moment it is live it refuses the
still-old-shape stored manifests, and composition freezes for every node that runs an affected
module.** Running services are untouched; only new declarations stop.
6. **Re-register each affected manifest under the new shape**, which the *new* controller accepts —
`module add <file> -source <repo> -ref <ref> -commit <commit>`. This writes the manifest to the
store without a build, so it is the fast way to lift the freeze. (The controller container is
distroless: `docker cp` the file to the container root `/x.json`; `/tmp` does not exist; the
root filesystem is writable. The file is lost when the container is recreated on the next image
swap, so copy it *after* the swap.)
7. `push --behind`, then verify `status` is clean and `seats` (or the relevant surface) shows the
new shape held by the right holders.
The freeze in a breaking change has been paid three times in one session (a fact-shape change, the
`/etc/hosts` region, a seat rename); each time it lasted seconds and no service dropped. It is
recoverable, but it is not something a push should trigger unwatched — which is the crux of what
automating this must solve.
## Why it is more than "add a webhook"
### 1. The trigger today is HAL's, not the mesh's
Build-on-push works for the catalogue because its repository has a Gitea webhook pointing at
`http://host.docker.internal:9877/webhook/gitea` — and **that receiver is `hal-gitea-tools.service`**
(`~/.hal/modules/hal/gitea/tools/server.js`), a *predecessor* component. The nox builder consumes
build work; it does not receive Git events. So the mesh's own build pipeline currently rides on a
HAL service, and:
- the `mesh-controller` repository was never wired to it, which is why the control plane is the one
thing that does **not** self-update — every controller deploy this session was `module moved` +
`build` by hand;
- when HAL is retired, build-on-push stops for the whole mesh.
**The mesh needs its own forge-webhook→build trigger**, a nox component (a module, and likely a
seat — `mesh-forge-trigger` or folded into the git seat's holder) that receives Git events and turns
them into build work over the broker, for **every** repository including `mesh-controller`. Replacing
`hal-gitea-tools` is the concrete first build. Its logic already exists to copy: match the pushed
repository (and changed paths, for a monorepo like the catalogue) against the build-context
repository of every registered module, and rebuild the matches.
### 2. The builder validates too — and a breaking change deadlocks it
The build machine embeds the same catalogue package the controller does, so **it validates a
manifest against its own compiled-in seat/schema set**. A breaking change therefore couples *four*
things, not two: the controller, the **builder**, every affected manifest, and every node's host.
This session's seat rename rebuilt the controller but not the builder, and the stale builder then
refused every manifest claiming a renamed seat.
Worse, one rename **deadlocked** the builder: the build machine's own seat was renamed
(`the-build-machine` → `mesh-build-machine`). To refresh the builder you must build it; to build it
the *running* (old) builder must accept the new builder's manifest — which claims the new name it
does not know. The old builder cannot build the new builder. Escapes:
- **Never rename a seat whose holder validates manifests** in an ordinary pass — the build machine's
seat belongs with the deferred delivering seats (ADR 0121). Reverting `mesh-build-machine` to
`the-build-machine` (deferred) lets the old builder build the new builder, which then knows the
new names.
- Or bootstrap a new builder image **out of band** (build locally, publish to the registry, register
the module at that digest), the way genesis loads the first builder — bypassing the old builder's
validation once.
Either way, self-update for breaking changes needs a **transition discipline** so a push does not
auto-freeze: the new control plane (and builder) should accept the *old and new* shape together for
one release — deprecated aliases in the seat set, a schema that reads both — then a later release
drops the old. With that, a breaking change rolls out on a push like any other: everything reads
both, the manifests migrate, the compatibility is removed. Without it, self-update would simply
automate the freeze.
## What to build
- **A nox forge-webhook trigger** (replaces `hal-gitea-tools`): receives Git events for every mesh
repository, dispatches build work to the builder over the broker, and records `module moved`
automatically. Wire `mesh-controller` to it so the control plane self-updates like everything else.
- **A transition discipline for breaking changes**: the control plane and builder accept old+new for
one release; the tooling that lands a schema/seat change emits the compatibility shim and the
follow-up that removes it. This is what makes step 4–7 above safe to trigger unwatched.
- **Config/package modules need no builder** — `module add` registers their manifest directly
(this is how the uplink managers and the re-registrations above were done). Only image-bearing
modules need the build machine, which narrows what the deadlock above can block.
## Why now, and why not yet
**Why it matters:** self-update is the difference between a mesh a person maintains by typing
pipeline steps and one that maintains itself, and it is a stated goal (parity with the predecessor's
pipelines). The HAL trigger dependency also makes it a retirement blocker: build-on-push dies with
HAL.
**Why not reflexively:** the trigger is a new component with the broker and forge in its blast
radius, and the transition discipline changes how every breaking change is written. Both should be
designed, not bolted on beside a freeze. The manual process above is the interim, and it works.
## References
- [ADR 0121](../../02-DECISIONS/0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md)
— the seat rename whose migration and builder deadlock this record is drawn from
- [ADR 0128](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md) — the
fact-shape change that first showed the breaking-change freeze
- `hal-gitea-tools.service` (`~/.hal/modules/hal/gitea/tools/server.js`) — the predecessor webhook
receiver on `:9877` the mesh currently rides on
- mesh-controller `cmd/mesh-builder` (the build machine), `internal/catalogue` (the seat/schema
validation the builder shares with the controller)
@@ -0,0 +1,65 @@
---
layer: to-be
status: proposed
code: []
updated: 2026-09-27
decisions:
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
---
# 31 — A module declares its fail2ban jail, and the mesh composes them per node
**A node's intrusion filter should be composed from the modules it runs, the same way its firewall
is.** The mesh already derives a node's nftables ruleset from every assigned module's `listens` and
`guards` (the `Filtering` mechanism). fail2ban is the same shape and is not modelled: a module that
runs an authenticating service — postgres, mssql, mailu — has a jail (a filter that reads its log
and a jail stanza that bans on it), and which jails a node's fail2ban runs should be exactly the
jails of the modules assigned to that node.
The predecessor did this with per-module files: `postgres` shipped `postgres-auth.conf`, `mssql`
shipped `mssql-auth.conf`, `mailu` shipped `mailu.conf`, and the node's fail2ban read whichever were
present. When HAL retired on novox those became dangling symlinks — fail2ban ran the jails only from
memory, and a restart would have dropped them. The base was salvaged (the fail2ban module now ships
`sshd`, `recidive`, and the `ignoreip` that spares the mesh's own range), but the **service jails
are gone**, because no nox module declares one yet.
## The shape
- **A module declares its jail in its manifest**, naming no node and no path (ADR 0112): the filter
(the failregex, or a stock filter it uses) and the jail stanza (port, logpath, maxretry, bantime).
The `postgres` module says what a postgres brute-force looks like and how to ban it; it does not
say on which machine, because it does not know.
- **The mesh composes them per node.** For each node, the jails of its assigned modules are gathered
and written into the fail2ban holder's `jail.d/` (and filters into `filter.d/`), exactly as
`listens`/`guards` are gathered into the node's firewall. So a node running postgres gets the
postgres jail; a node not running it does not. The `node-intrusion-prevention` holder receives
them the way a provider receives its consumers' contributions.
- **The base stays the fail2ban module's**: `sshd`, `recidive`, and the `ignoreip` naming
`${machine:mesh-range}` so a tunnel peer is never banned.
## Why this, and not the module writing the file itself
A module could declare a `file` resource at `/etc/fail2ban/jail.d/<x>.conf` directly. Rejected: the
path is the fail2ban holder's to own (one module owns `jail.d`, as one module owns the firewall
table), the jail's logpath and defaults want the mesh's composition (the `ignoreip`, the ban action
the node uses), and two modules writing into one directory is the collision the seat/holder model
exists to prevent. The module declares *what its jail is*; the holder's composition decides *how it
lands* — the same split as `listens` (the module says the port; the mesh says the rule).
## Why now
fail2ban on novox currently runs the service jails from memory only; the next restart drops them
(the `ignoreip` is safe on disk, so the mesh-partition risk is closed, but postgres/mssql/mailu
auth-banning would be lost). This is the mechanism that restores them properly, and it is needed as
each of those modules migrates to the other nodes — ace running postgres should get the postgres
jail, composed from the postgres module's manifest, without anyone editing a node.
## References
- [ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) — a module
names no node or path; its jail is declared the same way its `listens` are
- mesh-controller `internal/catalogue/adoption.go` (`Filtering` — the firewall composition this
mirrors), `internal/catalogue/manifest.go` (`Listens`/`Guards`, the fields a jail field sits
beside)
- mesh-catalog `modules/fail2ban` (the base: sshd, recidive, ignoreip); the service modules
(`postgres`, `mssql`, `mailu`) that will declare jails
@@ -0,0 +1,502 @@
---
layer: to-be
status: in-progress
code:
- mesh-controller internal/catalogue/declaration.go
- mesh-controller internal/catalogue/manifest.go
- mesh-controller internal/link/serve.go
- mesh-controller internal/link/bus.go
- mesh-controller internal/broker/nats.go
- mesh-controller internal/inventory/nodes.go
- mesh-host internal/apply/apply.go
- mesh-tools src/main.ts
- mesh-catalog modules/mesh-catalog
updated: 2026-09-28
decisions:
- 02-DECISIONS/0126-a-module-declares-its-own-seats.md
- 02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md
- 02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md
- 02-DECISIONS/0106-the-bus-is-nats.md
- 02-DECISIONS/0041-events-are-a-relationship.md
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
- 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
- 02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md
- 02-DECISIONS/0135-a-module-version-prepares-its-state-before-it-runs.md
- 02-DECISIONS/0136-a-step-gates-its-module-not-the-machine.md
- 02-DECISIONS/0134-the-mesh-says-what-it-applied.md
---
# 32. What a module declares, and what the bus makes of it
**A module that speaks to the mesh requires the bus, and receives what it needs to connect**
([ADR 0128](../../02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md)). What a module
declares are *relationships*; subjects, streams, consumers and permissions are all derived from
those, and a manifest never contains one.
> **Revised 2026-09-26.** This document opened by calling the bus *ambient* — "no module requires
> it, the way no module requires a filesystem". Two counts say otherwise: of 72 modules in the
> catalogue, **49 take a broker credential and 23 do not**, so an ambient connection would mint an
> account for a third of the catalogue that never speaks; and the 49 each hand-write the path it
> lands at, which is provisioning done badly by hand. The bus is required, and a module that does
> not require it has no account at all.
**The requirement delivers the connection; the declarations shape the authority.** `requires:
mesh-bus` says *this module talks to the mesh* and grants no subject by itself. `emits`,
`consumes`, `tools`, `uses` and a declared seat say what it may say and hear. Declaring a subject
without requiring the bus is incoherent and refused at registration.
This document is the declaration model. [Design 25](25-the-bus-on-nats.md) is the bus itself —
subjects, streams, accounts, enrolment — and stays the authority on the wire.
[Design 19](19-the-module-protocol.md) is the specification an SDK implements, and is rewritten
onto this in step 3 of [ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md).
## 1. A module names locally; the mesh derives the subject
This is the load-bearing rule.
[ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) says a module
definition names no node, mesh or path. A transport address is the same class of thing: if
manifests held literal subjects, reorganising the subject space would mean editing every module in
the catalogue, and the mesh would have hundreds of copies of a decision it made once.
| declared | derived |
|---|---|
| `emits: order.placed` | publish on `mesh.mod.<module>.event.order.placed` |
| `consumes: billing.order.placed` | durable consumer on `mesh.mod.billing.event.order.placed` |
| `tools: status` | queue-group subscription on `mesh.mod.<module>.tool.status` |
| seat `telegram-sender`, `accepts: send` | work-queue consumer on `mesh.seat.telegram-sender.accept.send` |
| `uses: telegram-sender` | publish on that seat's `accept` subjects, and nothing else |
**Wildcards, and they are the mesh's rather than a bus's.** *Added 2026-09-27, from
[issue 127](../../04-ISSUES/127-a-module-event-derives-a-subject-nothing-publishes/00-report.md).* A
consumer may write `*` for one name and `**` for the rest: `*.download.completed` is that event from
any module, and `**` on its own is every event in the mesh, which an audit logger wants and says in
one token. Spelled this way rather than the wire's, for the reason everything else here is local —
the bus the mesh runs on today spells these `*` and `#`, the one being built spells them `*` and
`>`, and a manifest naming either would stop being true when the wire changed. An emitted event
carries no wildcard: it names one event.
**A module publishes under its own name, and an event about a role belongs to the seat.** *Added
2026-09-27, same source.* The bus enforces that a namespace belongs to the module it is named for, so
an event named for somebody else cannot be published at all. Where the event is really about a role —
"the artifact store accepted an image" — the seat is the right home, because that name outlives
whoever fills it, and a consumer written against the holder's own name breaks when the holder
changes. **Not yet possible in practice**: seats carry protocol in the manifest and in the permission
model, and the shared library has no way for a module to publish on one. Until it does, such an event
lives under the emitting module's own name and the consumer carries that coupling.
**How the rule is checked, because it was not.** *Added 2026-09-27, same source.* Two checks, because
the mistake happens at two scales. Per manifest, at registration: an event is a local name, and the
old bus's form is refused with the name to write instead. Across the whole catalogue, as a test:
where a consumed event's emitter is present, it must emit that event. The second cannot demand a live
emitter for everything — a module lives in its own repository and may be installed long before the
one whose events it wants — so it says nothing about an absent emitter and everything about a present
one. **A subscription that matches nothing is not an error, it is silence**, which is why nothing
reported thirty-seven manifests being wrong the same way.
**It is `tools:`, not `serves:`.** Revision, found while implementing: the manifest already uses
`serves` for the facts a consumer needs in order to reach a provision, and two meanings under one
key in the file a module author reads most is a footgun. Worth noting that until now a module's
tools were not declared at all — they were known only at runtime, from an environment variable in
its image — so declaring them is new, and is what lets the mesh check that a module claiming a
seat answers what that seat's protocol promises.
**The `event` / `tool` / `accept` token is load-bearing, not decoration.** Revision, found while
defining the streams: a stream is defined by a subject filter, so a namespace holding both a
module's events and its tool calls cannot be filtered into an events stream without capturing
every tool invocation in the mesh — and a tool call must never be persisted
([design 25](25-the-bus-on-nats.md) §3 keeps tools on core NATS, where a lost call is a timeout the
caller already handles). The kind token is what makes `mesh.mod.*.event.>` a safe filter. The
first draft of this table had no token, which reads better and cannot be implemented.
**The test this must pass: the manifest survives the wire changing.** Reorganise the subject space
and every manifest in the catalogue is still correct. That is the property
[ADR 0039](../../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md) gave the sdk, applied to
declarations.
**A handler names an event the way its manifest does.** *Added 2026-09-27, found while fixing issue
127.* The runtime handed a handler the event name alone, so a manifest declaring
`consumes: builder.built` produced a pattern that could never match the key it was compared against,
and a module consuming one event from two emitters could tell them apart only by reading a header.
The subject already carries the emitter, so the key a module sees names it too — which makes a
disagreement between a manifest and the code a typo rather than a category error.
## 2. Three namespaces, and nothing else
**Its own** — `mesh.mod.<module>.>`. Its events and its tools. Nothing else may publish into it,
so an event's source is a fact the bus enforces rather than a claim in the body.
**Seats it holds** — `mesh.seat.<seat>.>`. Full participation: consume what the seat accepts,
publish what it emits, serve what it serves.
**Seats it uses** — publish only, and only on the `accepts` half. A sender cannot subscribe to a
seat's inbound subject and watch other modules' traffic, and cannot publish the seat's outbound
events and lie about outcomes.
A module naming anything outside these three is refused at registration. The whole permission set
is derivable from the declaration; nobody writes an access rule.
## 3. Queues are derived, never declared
A module says what it reacts to, not how delivery works. Each `consumes` becomes one durable
consumer; a seat's `accepts` becomes one work-queue consumer with a queue group named for the
seat. The module does not name them, does not know their names, and cannot misconfigure them —
and the controller stays the only writer of stream and consumer definitions
([design 25](25-the-bus-on-nats.md) §3).
**The mesh's own seats carry protocol too.** *Added 2026-09-27,
[ADR 0129](../../02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md).* A seat declared by a
module says what it accepts, emits and serves; the `mesh-*` set said only who does a job. So the mesh
had roles it could not describe — a build machine with three audiences for one outcome and no way to
derive a grant for any of them, and an event genuinely about a role with nowhere to live but the
namespace of whichever module happens to hold it. The mesh's seats now take the same three fields, and
the same machinery derives the holder's authority, its work queue and its consumers.
**So a build is work submitted to a role, like any other.** The build machine seat accepts a build and
emits an outcome, and the dedicated branch that carried builds retires: a work queue shared by several
machines is exactly what `accepts` already is, and a second mechanism for it is two places a permission
can be wrong.
**And one publish reaches three audiences without anybody's inbox being opened.** A build's outcome is
the seat's own event: whoever asked matches it by the id their request carried, the controller records
it, the catalogue places it in the graph. That is the fan-out a shared exchange gave for free, written
as a subject the mesh derived instead of a topology somebody configured — and it is why a holder needs
no permission to publish into an asker's inbox, which is the one grant design 25 §4 refuses by name.
**Retention belongs to whoever owns the namespace, not to a consumer.** A seat declares how long
its inbound backlog survives, because that is a property of the service:
```
seat: telegram-sender
scope: mesh
accepts: send retain 7d
emits: delivered, failed
serves: status
```
If each consumer could tune it, the mesh's durability would be an emergent property of whichever
manifest was edited last.
**Why a module's own events do not carry their own retention, though the same rule would allow
it.** A seat owns its namespace and gets a stream of its own, so it can say. A module's events
share one `EVENTS` stream, and three facts about JetStream decide that they must:
- **Storage is not a property of a subject.** A subject is only an address; a *stream* is a
separate object that captures subjects matching a filter. So "this topic is durable" is always
really "some stream covers it", and something has to create that stream.
- **Overlapping streams are refused, not merged.** Verified against the server: a per-module
stream beside a shared `mesh.mod.*.event.>` is rejected with *subjects overlap with an existing
stream*. So "one stream by default, its own for a module that wants different retention" is not
available — it is all of one or all of the other, and a filter cannot express an exception
either.
- **A stream per module breaks cross-module consumption.** An audit logger consuming every
module's events is one consumer on one stream today; with a stream each it becomes one consumer
per module, created and destroyed as modules come and go.
So: one stream, and **per-subject caps** for the fairness that actually matters — a noisy emitter
cannot evict a quiet one, which is verified (a cap of three, ten messages on one subject and one
on another, leaves four). What is genuinely unavailable is a different *age* per module, because
JetStream ages per stream and not per subject. A module that truly needs its own retention has a
way to say so: declare a seat, which owns its namespace and gets its own stream.
**Scope gives per-node workers without a new concept.** A module running on three nodes that each
need their own queue declares a node-scoped seat: one holder per node, three queues, same
machinery. Mesh-scoped and node-scoped seats already exist; here they do the work of "one shared
service" versus "one worker per machine".
## 4. Five relationships
| | provision | event | job | state | tool |
|---|---|---|---|---|---|
| shape | 1:1 resource | 1:many | N:1 | 1:1 | 1:1 |
| addressed to | a provider | the emitter's own namespace | a **seat** | one node | a module or seat |
| who must act | the provider | nobody | exactly one holder | that node | the server |
| credential | sealed, per consumer | none | none | none | none |
| reply | — | none | none, or an event later | a report | awaited |
| retention | — | age and size | work queue, explicit ack | **last per subject** | none |
| declared | `provides`/`requires` | `emits`/`consumes` | seat `accepts` / `uses` | the mesh's own | `serves` |
**Job** is the one [ADR 0041](../../02-DECISIONS/0041-events-are-a-relationship.md) had no room
for. Its table has events at 1:many and provisions at 1:1; a module submitting work to a service
is neither. It is not an event, because an event is a broadcast nobody is obliged to act on and a
second holder would do the work twice. It is not a provision, because there is no resource and no
credential. What makes it safe is not cleverness in the subscribe call but the seat: exactly one
holder, so exactly one worker, by construction.
**State** is the shape the deploy path needs and nothing else uses. A declaration is not an event
— replaying yesterday's is actively harmful — and not a job. Only the newest matters, which is
last-per-subject retention, and a node that has seen sequence *n* refuses *n−1* by construction.
That is the wire-level answer to
[issue 107](../../04-ISSUES/107-a-declaration-carries-no-order/00-report.md).
## 5. Seats
A module declares a seat with its protocol, and the mesh enforces one holder at its scope
([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md)). A caller declares that
it uses the *seat*, never the module, so the implementation can be replaced under it.
- The set of seats is **derived** — the mesh's own, plus every registered module's — so it is both
closed and extensible, and enumerating it is a query rather than an inventory.
- `mesh-*` is **reserved**: the prefix is the reservation rule, and a module declaring one is
refused at registration.
- Two modules declaring the same name: the second is refused.
- A module may not claim a seat whose protocol it does not implement.
- **Nobody holding a seat is not an error.** The stream exists from registration, so work queues
until a holder appears. Install the telegram module a week later and the backlog flushes.
## 6. The lifecycle: build, publish, deploy
Every shape above appears once, in order, and no step knows where the next one runs.
**A change lands.** The module holding `mesh-git` emits `pushed` — repository, ref, commit. An
event, because it is a fact about git and git's identity is the meaning.
**The change becomes work.** The controller consumes `pushed`, asks the catalogue which modules
are built from that repository and path
([ADR 0069](../../02-DECISIONS/0069-a-module-is-a-repository-and-a-path.md)), and submits one
**job** per affected module to the `mesh-build-machine` seat. The builder stays simple: it builds
what it is handed, and never resolves anything. A builder that dies mid-build has its job
redelivered, because a work queue with explicit ack is what that means.
**The artifact is published.** The builder pushes to the registry seats and emits `built` —
module, version, digest. An event again: a fact about the builder.
**The build cascade is that event fanning out through a graph the mesh already has.** A module
whose image is built *on* another's artifact declares that in `build.on`. So `built` reaches the
controller, which walks the declared graph and submits rebuild jobs for everything downstream. A
dependency cascade is not special machinery — it is one event, one derived graph, and the same job
queue.
**Deployment is state, not a message.** The controller composes each affected node's declaration
and publishes it last-per-subject. A node that was away gets exactly the current one, never a
queue of superseded ones, and a replayed older one is refused by sequence.
**A version prepares its state before it runs.** *Built 2026-09-28.* A module version may declare an
entrypoint that brings its state to the shape that version needs — the same vocabulary as the entrypoints it declares for its
tools and its provisioner, and nothing about how a machine runs it. The mesh runs that entrypoint as it
runs the module's own code, to completion, in the module's own context, and a version whose preparation
did not succeed does not run: the step gates that module and nothing else on the machine
([ADR 0136](../../02-DECISIONS/0136-a-step-gates-its-module-not-the-machine.md)), and the rollout stops
at the first machine that did not take it
([ADR 0135](../../02-DECISIONS/0135-a-module-version-prepares-its-state-before-it-runs.md), superseding
[ADR 0133](../../02-DECISIONS/0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md)).
Once per state, and the mesh derives what a state is: a consumer is a module on a machine, so what the
mesh provisions is per consumer and preparation is too. No level to choose, and no race to lock against.
**Applying is reported to a role.** The host applies and reports to the `mesh-controller` seat —
not to an address it was given at genesis. Held and retried while the store restarts
([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)).
**And the mesh says what it applied.** *Built 2026-09-28.* A report is control traffic only the control plane reads, so the
chain above went dark at the moment it touched a machine: nothing said which version a machine now runs,
or that it refused to. The control plane states those as facts under its own seat's namespace, when what
a machine runs changes rather than on every convergence pass, and anything that cares subscribes the way
the catalogue subscribes to `built` ([ADR 0134](../../02-DECISIONS/0134-the-mesh-says-what-it-applied.md)).
The facts are second-hand by design — one emitter, one ordering — and a machine that cannot reach the bus
produces none, so absence is not health.
What disappears across that chain is every address. No webhook URL, no registered callback, no
"which node is the builder on", no controller endpoint baked into a joining node. That is the
class of bug
[issue 102](../../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)
names, dissolved rather than fixed.
## 7. Modules depending on each other
Three kinds, and conflating them is how deployment ordering goes wrong.
**Build-time** — A's image is built on B's artifact. Resolved by the cascade above; nothing at
runtime cares.
**Provision** — A requires a database from B. A **hard** dependency: the credential must exist
before A can start, so resolution gates delivery and A is shown as waiting until B has answered
([design 27](27-a-module-requires-the-mesh-resolves.md)).
**Seat** — A uses B's seat. A **soft** dependency, and this is the one the bus changes. A starts
whether or not anyone holds the seat, because the stream absorbs the gap. Deployment order stops
mattering for everything expressed this way, and a service being restarted, moved or upgraded is
not an outage for its callers — it is latency.
That difference is worth choosing on purpose. A dependency expressed as a provision must be
ordered; the same dependency expressed as a seat need not be.
## 8. Versioning a protocol
A seat's protocol is a compatibility surface between modules that do not know each other and are
deployed at different times. Four ways it can change, and they are not equally dangerous:
| change | example | detectable |
|---|---|---|
| **additive** | a new `accepts` subject, a new optional field | nothing breaks |
| **removal or rename** | `send` becomes `deliver` | yes, mechanically |
| **shape** | an optional field becomes required | yes, if shapes are specified |
| **semantic** | `send` starts meaning *queue for tomorrow* | **no** |
**Additive is free.** A seat may grow without a version, without re-registering a caller, and
without ceremony. Most change is this.
**A breaking change is refused while anyone is bound.** Registration computes a compatibility
fingerprint over the seat's protocol — its subjects and the shapes they carry. A registration that
alters the fingerprint while callers are bound is refused, **and the refusal names them**. The
mesh already holds the `uses` graph, so this is derived rather than declared, and it turns a
runtime breakage into a registration-time conversation.
**When a break is genuinely needed, the version goes in the subject, not the name.** The seat stays
one thing; `mesh.seat.<seat>.v2.<verb>` runs beside v1 and the holder serves both. A caller moves
when it is ready. Versioning the *seat name* was considered and rejected: it forks the role, so
"one holder" stops meaning one provider of the capability, and every document naming the seat has
to be found and changed.
**Binding is recorded, not inferred.** A caller declares `uses: telegram-sender` with no version,
and resolution binds it to the current one and records that — the same **pin** machinery
[design 27](27-a-module-requires-the-mesh-resolves.md) already uses when resolution had a choice
to make. Moving to v2 is a deliberate re-pin, so nothing drifts onto a new protocol because it
happened to be newest.
**Retirement is reported, never automatic.** When the `uses` graph shows nothing bound to v1, the
overview says it is retirable. The mesh does not remove it.
**And none of this catches a semantic change.** Same subject, same shape, new meaning: no
fingerprint sees it, and no check proposed here would. The defences are review, and pushing
meaning into shape wherever it can go — a required `channel` field is caught, a changed
interpretation of an existing one is not. Saying so is better than implying the fingerprint is
complete, because a team that believes it is complete stops reviewing for the case it misses.
## 9. Provisioning over the bus
Provisioning rides the bus, and the provider stops having an address.
| part of a provision | shape |
|---|---|
| the requirement resolving to a provider | the controller's, not the bus's |
| the grant reaching the provisioner | request/reply to a **role** |
| `holds` — the reconcile question, every minute | the same call, on a timer |
| `provisioned` / `deprovisioned` | events |
| the credential reaching the consumer | §10 — fetched, never carried |
A provisioner's interface is already three calls — create, remove, holds — which is exactly a
`serves` protocol. So **a provision interface is a seat whose protocol is those three**, which is
why [design 26](26-the-seats.md) already allows a seat to deliver a provision. The two concepts
were converging before this document; here they meet.
What stays different, and must not be unified away: a provision has a **per-consumer resource and
a sealed credential**, created and destroyed per consumer. A seat protocol has neither — it is a
role you send to. Collapsing them would mean pretending a database is a subject.
### `mesh-bus` and `nats` are two interfaces, never one name
The mesh's own bus is **`mesh-bus`**, delivered by the `mesh-broker` seat and answered by the
controller — because the bus's accounts are configuration rather than something a provisioner
creates, so there is no provisioner process in the path and nothing waiting on a bus account in
order to make bus accounts. A module that runs a NATS server of its own and offers it as a
backing service provides **`nats`**, exactly as the deprecated broker provides `amqp`
([ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md): AMQP is not a provision, and everything speaks to the `mesh-broker` seat)).
They are never the same name. A manifest saying `nats` could otherwise mean either the mesh's
nervous system or a private queue, and the difference between those is the whole architecture.
0119's rule decides which is legitimate: a private bus is a backing service, never a channel to
another module.
### Where addresses survive
"Where is it?" is two different problems, and the bus solves one of them completely and the other
not at all. Keeping them apart matters, because a reader who thinks the mesh no longer has
addresses will believe a class of bug is fixed when it is untouched.
**Something that is on the bus: the address disappears.** A host reporting in used to need the
controller's address — recorded somewhere, at some moment, and wrong as soon as anything moved.
Now it publishes to `mesh.seat.mesh-controller.report` and the bus routes it to whoever holds the
seat. Nothing anywhere records where the controller is, so nothing can record it *wrongly*. The
same is true of the builder, the catalogue, the telegram sender. This class is not mitigated; it
is gone, because the information is no longer stored.
**Something that is not on the bus: the address stays, exactly as before.** A module that requires
a database does not reach postgres over NATS — it opens a postgres connection, because postgres
speaks postgres and is not listening on any subject. Its credential contains a host and a port,
and no amount of subject addressing changes that.
So the fix for that second class is unchanged and is not this document's:
[ADR 0098](../../02-DECISIONS/0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md) —
fetch the fact where it is used rather than storing a copy — which is what
[issue 102](../../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)
is actually about. The bus makes that class *smaller* by removing every mesh-internal address from
it. It does not make it empty.
**And one address is irreducible: the bus's own.** A node has to know where the broker is before
it can use subjects for anything, so that one cannot be a subject. It is the addressing equivalent
of §10's bootstrap — the first thing cannot be found by the mechanism that finds everything else.
## 10. Secrets, and why they never enter a stream
Everything the vault does is request/reply to the `mesh-vault` role: mint, fetch, rotate. In that
sense it is as much on the bus as anything else.
**The bus is not trusted with a secret, and does not need to be.** A secret is sealed to its
recipient, so what crosses the bus is ciphertext only that recipient can open. The broker sees
that a secret moved, and to whom — metadata, which is acceptable — and never a plaintext.
**But sealed is not enough on its own, because a stream persists.** A sealed secret written into
a JetStream stream is a durable ciphertext sitting in the mesh's own storage, and the day a
sealing key leaks, that stream is an archive rather than a moment. So:
- **A secret travels on core request/reply, never through a stream.** No persistence, no replay,
nothing to exfiltrate later.
- **A declaration names a secret; it does not carry one.** Declarations are the state shape, which
*is* a stream — so the host fetches the secret from the vault at apply time, over the core path.
That is [ADR 0098](../../02-DECISIONS/0098-a-fact-a-provider-makes-at-first-start-is-fetched-from-it.md)'s
existing discipline — *fetched from it, not carried* — applied to the one payload where carrying
it is worst.
**The bootstrap, which is circular and has a precedent.** The vault makes every secret
([ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md)), including the bus's own
passwords. The vault is a module, and a module needs a bus account, whose password the vault
makes. Nothing can go first.
This is the shape [ADR 0067](../../02-DECISIONS/0067-genesis-is-a-pivot.md) already resolves for
the control plane: **genesis is a pivot.** The controller mints the handful of foundation
credentials itself, raises the store, the broker and the vault, and then the vault takes over and
mints everything from there — the same move as raising a temporary control plane and reinstalling
it as an ordinary module once the registry exists.
So there are exactly two things the normal path cannot make, both at genesis, both ending the
moment the mesh can mint for itself: **the bus's own accounts** (§the bootstrap argument in
[ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md): AMQP is not a provision, and everything speaks to the `mesh-broker` seat) (superseding [ADR 0125](../../02-DECISIONS/0125-the-bus-is-the-only-broker.md)) — a provisioner is a module and
needs an account before it can run) and **the vault's own credential**. Any third exception is a
design failure, and naming these two is what makes a third one visible.
## 11. Open
**Semantic change has no mechanical defence** (§8). Recorded as open rather than solved, because
it is the residue of a question the rest of §8 answers and the part a fingerprint cannot reach.
**Whether a module may declare a seat it does not itself claim** — the contract as one thing, the
implementation as another, which is how two competing implementations would ever exist.
**Whether a container should have a readiness notion.** Only an action carries `verify`, so a step that
must run once a service *answers* — seeding through its own API — cannot be declared at all. Named here
because the steps above make the gap obvious, not because they caused it.
**Whether `consumes` naming another module couples too tightly.** It is kept here deliberately —
an event's provenance is its meaning — but a consumer of `billing.order.placed` does depend on
billing existing under that name.
## 12. How it is checked
- **A manifest holds no subject.** A catalogue test: no manifest contains a string matching the
subject grammar. The rule is worthless if it is followed by convention.
- **A preparation is given what the module is given.** A composition test: what the preparation
entrypoint receives equals what the module's own code receives, asserted rather than written twice —
which is the drift a hand-written step invites, three times over in the catalogue today.
- **A convergence that changed nothing says nothing.** Two identical reports, one emitted fact: what is
guarded against is a fact per minute per machine, which is a stream nobody reads.
- **Permissions are exactly the three namespaces.** A composition test per module: the derived
permission set equals what its declaration implies, and a hand-written addition to it fails.
- **A sender cannot read the queue it writes to.** A bed: a module declaring `uses` is refused
subscribe on that seat's inbound subject.
- **One holder, one delivery.** A bed: a seat's job delivered once with the holder running, and
a second claim of the seat refused.
- **A queued job survives no holder.** A bed: submit with the seat unheld, assign the holder,
the job is delivered.
- **The cascade rebuilds exactly the dependents.** A bed: publish an artifact two modules build
on, and exactly those two are rebuilt.
- **A stale declaration is refused.** A bed: replay sequence *n−1* after *n*, and the node refuses
it rather than applying it.
@@ -0,0 +1,137 @@
---
layer: to-be
status: designed
code: []
updated: 2026-09-28
decisions:
- 02-DECISIONS/0132-a-seat-carries-the-tools-its-holder-must-serve.md
- 02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md
- 02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md
---
# 33 — The tools the mesh answers
**An agent can call the mesh's tools and cannot find out what they are.** Both halves were measured
on the live mesh on 2026-09-28: a client holding an operator credential connected to the bus, asked a
module for its repositories and got them; the same client's request for the tool list found nothing
serving it. The transport works, the account model works, the adapter that speaks the agent protocol
works. What is missing is the mesh being able to say what it can do.
This design is the answer to that question, and it has three families in it, because a tool belongs to
whoever is accountable for answering it.
## 1. Three families, and why the split is not arbitrary
| Family | Addressed to | Where the definition lives | Example |
|---|---|---|---|
| A **role's** tools | the seat: `mesh.seat.<seat>.tool.<verb>` | the seat's protocol, in the mesh's records | ask *the forge* to list its repositories |
| A **module's** tools | the module: `mesh.mod.<module>.tool.<name>` | that module's code | ask *this gitea* for `gitea_list_repos` |
| The **mesh's** own verbs | the `mesh-controller` seat | the seat's protocol, as above | `status`, `push`, `build`, `assign` |
The split follows accountability. A role is something the mesh guarantees exactly one holder of, so
what the role answers is the mesh's to define and a holder's to implement
([ADR 0132](../../02-DECISIONS/0132-a-seat-carries-the-tools-its-holder-must-serve.md)). A module's
own tools are nobody's business but the module's, and their definitions live where they are
implemented, because a copy kept anywhere else drifts from the code that answers.
The mesh's own verbs are the third family only in where they come from, not in kind: the control plane
holds a seat like anything else, and its tools are that seat's. This is what keeps them addressable
while the control plane is being replaced, which is the moment they are most needed.
**Both names for one capability is deliberate and bounded to this.** A forge holding the `git` seat
answers the role's `list_repos` and its own `gitea_list_repos`, because the same module may run
without the seat — a second instance, kept for one purpose — and then only the second name is true.
The caller chooses which question it is asking. Nothing else in the mesh gets two names.
## 2. What a seat's tool is
A verb, what it does, and the schema of its arguments and its answer. A name alone is not callable by
something that has never seen the mesh before, which is the whole population this surface exists for.
The protocol a seat carries today is three lists of bare verbs
([ADR 0129](../../02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md)), and it must widen to
carry the rest. Two constraints on that widening:
- **It lives in the mesh's records, not in the control plane's binary.** Today a seat's protocol comes
from compiled defaults, merged in as a row is read, because the seat rows never gained the columns.
Discovery that reads a binary is discovery that disagrees with the mesh the moment the two are on
different versions.
- **The schema is stated in the form an agent protocol already uses**, so nothing translates between a
seat's idea of an argument and the caller's. A translation layer would be a second definition of
what a tool is.
## 3. Holding a seat means serving its tools
A module may not occupy a seat unless it serves every verb that seat declares. This joins the
conditions of holding that already exist — providing what the seat delivers, being assigned at the
seat's scope — and is refused the same way: at registration and at handover, naming the verbs that are
missing rather than the fact that something is.
A module knows which seats it claims, so knowing which tools it must serve is not a discovery problem
for the module: the seat says, the module implements, and anything beyond that is its own.
## 4. Addressing a node-scoped seat
A seat's subject is flat today — `mesh.seat.<seat>.<kind>.<verb>` — which is correct for a seat the
mesh has one holder of and wrong for the six node-scoped seats, where one subject would reach every
machine's holder and the holders' queue group would hand the call to whichever answered first. A
node-scoped seat's tool therefore carries the node it is asked of. Nothing about a mesh-scoped seat
changes.
## 5. Discovery
**What a role answers is a read.** The seats and their protocols are records, so the list is a query
against the mesh's own store: no call to a module in the path, nothing that has to be running, and an
answer that stays true while a holder is restarting or being replaced.
**What a module answers comes from the module.** Its definitions live in its code, so it is asked, and
the answer is as available as the module is — which is the right coupling for a tool that only exists
while that module does.
A caller therefore gets one list assembled from two sources, and the difference is visible in it: a
role's tool names a seat, a module's names a module. An agent that wants to survive a holder being
replaced binds to the first.
## 6. What serves this to an agent
A module the mesh assigns to the machine where the agent runs, holding a credential the mesh minted,
with authority derived from what it may call — not a program started by hand with a credential printed
to a terminal. The adapter itself already exists and is thin by design; what changes is that it stops
being something a person carries and becomes something the mesh runs, on a node, like everything else.
An agent's authority can then be role-shaped: *the forge's tools*, rather than a list of
module-specific names that changes the day the forge is replaced.
## 7. Versioning
A seat's tools are an interface and change like one. Additive within a version. A change that would
break a caller takes the version token the subject already has room for, and the two versions run side
by side until nothing is bound to the old one.
## How it is checked
- **A holder missing a verb cannot take the seat.** One test per condition of holding, as the existing
conditions have, and the live refusal names the verbs.
- **A verb nobody declared is a subject nobody may use.** The bus grants are derived from the seat's
protocol already, and the golden composition of the user list is what keeps that honest: a holder is
granted exactly the seat's verbs, a user of the seat exactly the publish side.
- **Discovery needs no running module.** The test for a role's tools reads records and asserts the
answer equals what the seats declare — if it needed a module up, it would not be a read.
- **Two nodes holding one node-scoped seat derive two addresses.** Checked by the same test as the rest
of the subject table.
## What this does not settle
- Which verbs each seat should serve. That is a decision per seat, and the reason to do it slowly: a
seat's tools bind every future holder.
- Whether a module's own tool definitions should also be recorded when a build resolves its manifest.
There is an argument for it — the mesh could then answer for a module that is down — and an argument
against, which is that a recorded copy of a live definition is a copy that can be wrong.
## References
- [ADR 0132](../../02-DECISIONS/0132-a-seat-carries-the-tools-its-holder-must-serve.md) — the decision this designs
- [ADR 0129](../../02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md) — a seat carries the protocol of its role
- [ADR 0095](../../02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md) — a tool call passes one process where an audit belongs
- [`26-the-seats.md`](26-the-seats.md) — what a seat is, how it is held and handed over
- [`25-the-bus-on-nats.md`](25-the-bus-on-nats.md) §7 — a person's account, their inbox, and the adapter
+4 -2
View File
@@ -35,11 +35,13 @@ document is written and this one's status becomes `implemented`.
| [`23-choosing-a-provider.md`](23-choosing-a-provider.md) | Which of several providers of a kind serves a consumer, and when a module carries its own instead | [ADR 0084](../../02-DECISIONS/0084-which-provider-serves-a-consumer.md), [ADR 0027](../../02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md) | | [`23-choosing-a-provider.md`](23-choosing-a-provider.md) | Which of several providers of a kind serves a consumer, and when a module carries its own instead | [ADR 0084](../../02-DECISIONS/0084-which-provider-serves-a-consumer.md), [ADR 0027](../../02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md) |
| [`24-the-secrets-vault.md`](24-the-secrets-vault.md) | The module that owns a secret — a `secret` provision, and the boundary of what it owns | [ADR 0085](../../02-DECISIONS/0085-a-secret-is-a-provision.md), [ADR 0031](../../02-DECISIONS/0031-the-control-plane-authenticates-nobody.md), [ADR 0048](../../02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md) | | [`24-the-secrets-vault.md`](24-the-secrets-vault.md) | The module that owns a secret — a `secret` provision, and the boundary of what it owns | [ADR 0085](../../02-DECISIONS/0085-a-secret-is-a-provision.md), [ADR 0031](../../02-DECISIONS/0031-the-control-plane-authenticates-nobody.md), [ADR 0048](../../02-DECISIONS/0048-a-provider-creates-the-credential-the-mesh-minted.md) |
| [`25-the-bus-on-nats.md`](25-the-bus-on-nats.md) | **Proposed.** The architecture of the mesh's bus on NATS: what rides which subject under which guarantee and whose account, how a node joins, how a person reaches a tool, and how the mesh moves from the bus it has | [ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md), [ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md), [ADR 0043](../../02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md), [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md) | | [`25-the-bus-on-nats.md`](25-the-bus-on-nats.md) | **Proposed.** The architecture of the mesh's bus on NATS: what rides which subject under which guarantee and whose account, how a node joins, how a person reaches a tool, and how the mesh moves from the bus it has | [ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md), [ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md), [ADR 0043](../../02-DECISIONS/0043-a-module-broker-account-is-scoped-by-emits-and-consumes.md), [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md) |
| [`26-the-seats.md`](26-the-seats.md) | **Proposed.** What a mesh can have one of, who fills each, and a seat's holder answering for the provision it delivers — including the `git` seat a build's source can live on | [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md), [ADR 0111](../../02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md), [ADR 0109](../../02-DECISIONS/0109-a-package-registry-seat-is-one-per-ecosystem.md) | | [`26-the-seats.md`](26-the-seats.md) | **Proposed.** What a mesh can have one of, who fills each, and a seat's holder answering for the provision it delivers — including the `git` seat a build's source can live on | [ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) (superseding [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md)), [ADR 0111](../../02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md), [ADR 0109](../../02-DECISIONS/0109-a-package-registry-seat-is-one-per-ecosystem.md) |
| [`27-a-module-requires-the-mesh-resolves.md`](27-a-module-requires-the-mesh-resolves.md) | **Proposed.** One concept for everything a module needs: a requirement with a contract, answered by one of four kinds of provider, resolved at assignment or refused. Retires settings, placeholders, facts and paths in definitions | [ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md), [ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md), [ADR 0114](../../02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md), [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md) | | [`27-a-module-requires-the-mesh-resolves.md`](27-a-module-requires-the-mesh-resolves.md) | **Proposed.** One concept for everything a module needs: a requirement with a contract, answered by one of four kinds of provider, resolved at assignment or refused. Retires settings, placeholders, facts and paths in definitions | [ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md), [ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md), [ADR 0114](../../02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md), [ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) (superseding [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md)) |
| [`28-building-the-bus.md`](28-building-the-bus.md) | **Proposed.** The five steps of the bus work in the order their dependencies allow, each ending at a bed — with the surface measured, so no step's size is a guess | [ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md), [ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md), [ADR 0074](../../02-DECISIONS/0074-the-wire-is-specified-not-the-types.md) | | [`28-building-the-bus.md`](28-building-the-bus.md) | **Proposed.** The five steps of the bus work in the order their dependencies allow, each ending at a bed — with the surface measured, so no step's size is a guess | [ADR 0116](../../02-DECISIONS/0116-the-bus-is-built-in-five-steps.md), [ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md), [ADR 0074](../../02-DECISIONS/0074-the-wire-is-specified-not-the-types.md) |
| [`29-a-node-has-operator-accounts.md`](29-a-node-has-operator-accounts.md) | **Proposed.** The mesh models machines but not the humans on them: a node gains operator accounts, and a resource may live under a home owned by its account — what would own ~/.ssh, dotfiles and ~/.config when HAL retires | [ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md), [ADR 0051](../../02-DECISIONS/0051-shared-data-is-the-operators.md) | | [`29-a-node-has-operator-accounts.md`](29-a-node-has-operator-accounts.md) | **Proposed.** The mesh models machines but not the humans on them: a node gains operator accounts, and a resource may live under a home owned by its account — what would own ~/.ssh, dotfiles and ~/.config when HAL retires | [ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md), [ADR 0051](../../02-DECISIONS/0051-shared-data-is-the-operators.md) |
| [`32-what-a-module-declares.md`](32-what-a-module-declares.md) | **Proposed.** What a module declares and what the bus derives from it: three namespaces, subjects from local names, queues never declared, the five relationships, and the build-publish-deploy lifecycle on one bus | [ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md), [ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md), superseded by [ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md) (superseding [ADR 0125](../../02-DECISIONS/0125-the-bus-is-the-only-broker.md)), [ADR 0041](../../02-DECISIONS/0041-events-are-a-relationship.md) |
## Not yet written ## Not yet written
- **The remaining six contexts.** - **The remaining six contexts.**
@@ -155,3 +155,17 @@ design document here, and get it back. That check fails today by design.
**What stands until then** is the signpost, and the honest description of it: reachable, not **What stands until then** is the signpost, and the honest description of it: reachable, not
surfacing. surfacing.
## Where this stands, 2026-09-29
*Added in a grooming pass.* The knowledge base this record is about is the **predecessor's**, and it
is no longer reachable from anything: the surface that answered `recall_search` speaks the transport
the mesh removed at the cut-over
([issue 147](../147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md)).
So the sentence in `README.md` that this record catches — *these documents are still indexed into
the knowledge base* — is now wrong twice over: nothing indexed them, and there is nothing to index
them into. The record stays open, and its answer is no longer "index this repository somewhere"; it
is whatever the mesh grows as its own knowledge surface, if it grows one. Until then the honest fix
is the README, which should stop claiming a property nothing provides.
@@ -1,8 +1,8 @@
--- ---
status: open status: resolved
opened: 2026-09-22 opened: 2026-09-22
located-in: [] located-in: [mesh-controller internal/link/protocol.go (the report field that was missing), internal/inventory, cmd/mesh-controller (node show and status)]
fixed-by: fixed-by: mesh-controller 7683ba8, corrected by 1d9c102
amended-design: amended-design:
--- ---
@@ -0,0 +1,143 @@
# 087 — resolved: the mesh knows which host runs a machine
*2026-09-30.*
## The field existed and was thrown away on arrival
The machine has reported its host version since
[ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) — `Host` on the report, with
a comment saying why it must be there: *"without it nothing can say a machine is behind."*
**The controller's own copy of the report did not have the field.** Two structs describe one message,
one on each side of the wire, and only the sending side had it — so it unmarshalled into nothing and the
mesh could not answer a question the machine had been answering for a week. That is the whole of this
issue's mechanism, and it is worth stating plainly because neither side was wrong on its own.
## What it says now
`node show` names it per machine:
```
last heard from here
host 2026-09-30-0214
```
`not reported — this machine has not said since the mesh began keeping it` where the mesh has not been
told, because a machine that has not said is a different thing from a machine running nothing.
`status` names the machines that are behind another:
```
1 machine(s) run an older host than another machine does:
ace 2026-09-29-0113
the newest any machine reports is 2026-09-30-0214. A host refuses a declaration carrying a
field it does not know, whole — so a new field reaches these machines last
```
## Disagreement, not staleness, and that is deliberate
The open questions asked whether the controller should refuse to send a declaration a node cannot
parse. It cannot yet, honestly: **nothing delivers a host version** (ADR 0141 is accepted and not
built, which is [issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md)),
so the mesh holds no canonical current version and "behind" has no fixed point to be behind.
What it can say truthfully is that these machines do not all run the same host, and which is newest of
the ones it has been told about. That is the fact that matters before a declaration gains a field: **the
oldest host in the mesh is what the mesh may send.**
Two deliberate refusals to guess:
- **A machine that has reported nothing is not called behind.** It may be running anything. `node show`
says it has not said, per machine, which is the honest form.
- **Versions compare as strings.** That suits the timestamps and commits this mesh uses and is wrong
for a scheme where `10` sorts before `9`. Said in the code at the place that would have to learn,
rather than left as a surprise.
## What I shipped first was wrong, and the mesh said so within the hour
The first version reported *"N machine(s) run an older host than another machine does"* and worked out
which by comparing versions as strings. **A host reports its version as a commit, and commits have no
order.**
On the live mesh, with the adopted machine pushed for the first time:
```
3 machine(s) run an older host than another machine does:
g14 04a27ca
novox 04a27ca
shanks 04a27ca
```
Those three run the **newer** host — installed 09:18, against the adopted machine's 01:13. `ced54d4`
sorts above `04a27ca` and that is all it means. An arbitrary lexicographic result, presented as a fact,
about the one thing this record exists to make trustworthy.
The code carried a caveat saying versions compare as strings and that this "is enough for the timestamps
and commits this mesh uses". That was the error, written down and not noticed: it is enough for
timestamps and it is **meaningless** for commits, and the mesh reports commits.
It now reports the split and claims no ordering:
```
4 machine(s) do not all run the same host:
04a27ca g14, novox, shanks
ced54d4 ace
a host refuses a declaration carrying a field it does not know, whole — so the mesh may
send only what every one of these understands. Which of them is newer is not readable
from a commit; that needs a version the host reports as ordered
```
More useful as well as more honest — the reader sees who is on which side of the split, which is what
decides whether a field can be sent — and it leaves the ordering where it belongs: with the host, which
would have to report something ordered for anybody to have it.
**This is [issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
arriving by my own door**, an hour after closing it: a report that confidently says the opposite of the
truth is worse than one that says less, because it trains a reader to distrust the whole surface.
## Measured on the mesh, and one limitation it exposed
All four machines run the identical host binary — same digest, installed within eighteen seconds of each
other — and at first only one reported its version. The other three said *not reported*, which read as a
difference between machines where there was none.
**A machine states its host version only when the mesh sends it a declaration.** The report is published
after an apply; the periodic reconcile that runs every five minutes publishes nothing, because it is the
machine keeping itself as declared rather than answering anything. So a machine that is current and idle
never says, and the mesh cannot distinguish that from a machine running something ancient.
Confirmed by pushing: before, `not reported`; after, `04a27ca` — the same version the machine that had
been pushed already reported.
```
novox host 04a27ca
shanks host 04a27ca
g14 host 04a27ca
ace host not reported — this machine has not said since the mesh began keeping it
```
`ace` has not been pushed since the field existed; it is adopted and parked.
**This is enough for the purpose and not enough for the claim.** For deciding whether a new declaration
field is safe it is sufficient, because pushing is what the mesh is about to do anyway and the answer
arrives with the act. For *knowing what the mesh is running*, it is not: a long-idle machine's entry is
as old as its last push, and the honest reading of `not reported` is "nobody has asked recently" rather
than "this machine is silent". The words say the first, which is why they are those words.
Making a heartbeat carry it would close the gap and is a change to what a heartbeat is — a bare word
that the node is there, deliberately carrying nothing else. Left alone rather than widened in passing.
## The open questions, answered as far as they can be
- *Should a node report the version of its host?* It already did. The gap was the reading.
- *Should the mesh refuse to send a field no node understands yet, or refuse per node and say so?*
Neither, yet — refusing needs the mesh to know which fields need which version, which is the third
question below and is not answered here. What it does is make the disagreement visible before
somebody adds a field.
- *Is there a general shape — a declaration saying which version of the host it needs?* Still open, and
now cheaper to answer: the versions are recorded, so a minimum-version field on a declaration has
something to compare against. It belongs with
[issue 107](../107-a-declaration-carries-no-order/00-report.md), which wants to add a field and is the
first thing this makes safe.
@@ -1,8 +1,8 @@
--- ---
status: open status: resolved
opened: 2026-09-22 opened: 2026-09-22
located-in: [] located-in: [mesh-catalog modules/gitea, mesh-controller internal/catalogue/declaration.go]
fixed-by: fixed-by: mesh-controller 7352c84, merged in #46 — a module is told its port in a container's environment too, as a file already was
amended-design: amended-design:
--- ---
@@ -39,3 +39,9 @@ the assignment happens to differ.
- Should composition refuse an environment value that names a port the module does not fix, the - Should composition refuse an environment value that names a port the module does not fix, the
way it refuses other claims a module cannot make? way it refuses other claims a module cannot make?
- Which other modules write their own address, with a port, into their environment? - Which other modules write their own address, with a port, into their environment?
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* The forge's address follows a moved port the same way every other reader does. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -1,8 +1,8 @@
--- ---
status: open status: resolved
opened: 2026-09-22 opened: 2026-09-22
located-in: [] located-in: [mesh-controller internal/catalogue/declaration.go, mesh-catalog (every routed module)]
fixed-by: fixed-by: mesh-controller bdf965d (a route names the endpoint it serves) with `portOfEndpoint` and `AtPublishedPort` — the contribution carries the endpoint's declared port and the machine-side redirection is applied to it like any other
amended-design: amended-design:
--- ---
@@ -45,3 +45,9 @@ precisely because the predecessor holds the usual one.
keeping the mapping out of rendered configuration? keeping the mapping out of rendered configuration?
- What should refuse a declaration whose contributed route names a port nothing on that node - What should refuse a declaration whose contributed route names a port nothing on that node
listens on? listens on?
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A route names an endpoint rather than a port, and the redirection that turns a declared port into the published one is applied to contributions too. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -1,8 +1,8 @@
--- ---
status: located status: resolved
opened: 2026-09-23 opened: 2026-09-23
located-in: [mesh-controller internal/link, mesh-host internal/link] located-in: [mesh-controller internal/link, mesh-host internal/link]
fixed-by: fixed-by: mesh-host PR 59 (the host refuses an older sequence and drains by it), mesh-controller PR 160 (each send is numbered under the node's hold) — measured 2026-09-30, 02-resolution.md
amended-design: amended-design:
--- ---
@@ -40,3 +40,16 @@ exists there and is thrown away at the wire.
separate genesis-digest branch is needed on the host? separate genesis-digest branch is needed on the host?
- Is a sequence enough, or does a mode change deserve its own marker, so a replayed converged - Is a sequence enough, or does a mode change deserve its own marker, so a replayed converged
declaration is refused by mode as well as by order? declaration is refused by mode as well as by order?
## What has since made this safer to do (2026-09-30)
Adding a `sequence` to a declaration is adding a field, and a host refuses a declaration carrying a
field it does not know — whole. That was
[issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md), and it is resolved: the
mesh now records which host each machine reports and `status` names every machine running an older one
than another does.
So the flag day is visible before it is walked into, which it was not when this was filed. It does not
make the field free: **the oldest host in the mesh is still what the mesh may send**, and one machine of
four is behind today. A sequence that an old host refuses takes that machine out of the mesh's reach
entirely — worse than the replay it prevents, which has never been observed.
@@ -0,0 +1,76 @@
# 107 — diagnosis: the fix is a flag day, and it should wait for delivery
*2026-09-30. Read, measured, and not built — deliberately.*
## The premise is confirmed
A host parses a declaration with unknown fields refused, and the code says why rather than leaving it
to be inferred:
> `DisallowUnknownFields` is the whole point rather than strictness for its own sake: a field the host
> does not know is a thing the control plane believes it asked for.
So adding `sequence` and `supersedes` is not an additive change. **Any host that has not been upgraded
refuses the whole declaration and applies nothing** — which is exactly the behaviour that keeps a
half-understood declaration off a machine, and exactly what makes this expensive.
## What has changed since this was filed
[Issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md) is resolved: the mesh now
records the host version each machine reports and `status` names every machine running an older host
than another does. The flag day is visible before it is walked into, which it was not on 2026-09-23.
That makes the cost measurable rather than hypothetical, and the measurement is the reason this is not
being built today.
## Why it waits
**One machine of four runs an older host, and it cannot be upgraded.** `ace` is adopted, deliberately
parked until the network and module-assignment work is settled, and **nothing delivers a host version at
all** — [ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) is accepted and not
built, which is [issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md).
Every machine takes a hand-placed binary.
So shipping the field means, in order: place a host by hand on three machines, unpark the fourth, place
it there too, and only then turn the controller half on. A machine missed in that sequence is a machine
the mesh cannot send anything to at all — not degraded, unreachable.
**And the fault it prevents has never been observed.** The record says so itself: *"Not observed;
constructed from the code, and narrow."* It needs a backlog of more than sixteen declarations queued
across a `converge`/`adopt` pair, or a broker slow enough to split one, and the host already applies the
newest of a drained batch and refuses a declaration that is not the last by digest.
**Trading a machine's reachability for a replay nobody has seen is the wrong way round.** The right
order is [issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md) first —
when the mesh can deliver a host, a declaration field costs a rollout instead of an expedition — and the
work order already puts that in its last group, as the proof that the mesh can make another of itself.
## What the open questions look like now
- *A per-node `sequence` under the controller's node hold, and `supersedes` as the previous digest?*
Still the right shape. The controller already holds the lock and already records each send, so the
order exists and is thrown away at the wire — unchanged since this was filed.
- *Genesis signing its bundle as sequence zero?* Yes, and it is the cheaper half: the bundle is written
by the host that will read it, so it has no flag day of its own.
- *Is a sequence enough, or does a mode change deserve its own marker?* A sequence alone does not stop
a replayed *converged* declaration reaching a node that has since been returned to adopted, which is
the incident of issue 104 by another door and is what this record names as its real risk. It wants
both, and the second is the one worth having first.
- **And one this record did not ask:** should a declaration say which host version it needs? 087 makes
that comparable for the first time, and it is the general form of the answer — a field that announces
its own requirement, rather than a flag day per field, for ever.
## Status
Left `located`. The owner is unchanged, the shape of the fix is agreed, and the gate is
[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md) rather than anything
in this record. **This is a judgement about order, not a refusal** — it is cheap to overrule, and the
code is a day's work once a host can be delivered.
## The gate has opened (2026-09-30, evening)
The mesh delivers the host now — built by its own toolchain, published to its own registry, delivered
over the bus and started by the launcher, on all four machines
([issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md)). A declaration
field is a build and a push, not an expedition. The order this record asked for — hosts first, then
the controller — is now two commands and a status line that says when the first has finished.
@@ -0,0 +1,60 @@
# 107 — resolved: a declaration carries its order
*2026-09-30. Measured on the mesh.*
## What was done
**Hosts first, then the controller** — the order [issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md)
says a new declaration field needs, and now a build and a push rather than an expedition
([issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md)).
The host understands a `sequence` on a declaration and tolerates its absence: absent reads as "no
order claimed", not "first", so a controller that sends none is still understood and a host that kept
a declaration before it understood the field compares nothing. It refuses a declaration with a lower
sequence than the one it kept, whole, and says why; and the drain that picks one declaration from a
batch keeps the highest sequence rather than the last to arrive — which is the case the report
constructed, a backlog drained out of order.
The controller numbers each send: the next number for that node, taken under the node's hold, before
the body exists, so the number is inside what the mesh signs and a replayed older declaration cannot
borrow a newer one's.
## Measured
```
push shanks; push shanks
sequence in kept declaration: 2
node sequence
novox 2
shanks 2
ace (none — not sent since numbering)
g14 (none)
status: nobody "not running what the mesh would send them"
```
Both applies went through; neither was refused; the machine holding the earlier one accepted the later.
## The subtlety, which would have read every machine as behind for ever
The mesh decides a machine is behind by comparing the digest of what it **would** send against what it
**did** send. A number changes the bytes. So the read-only comparison composes with the number the
machine was *last* sent — not a fresh one — and is byte for byte what was sent when nothing else
changed. Without that, numbering would have made `status` name all four machines as out of date on
every reading, permanently.
## The open questions
- *A per-node `sequence` under the controller's node hold?* Yes, as described. **`supersedes` — the
previous digest — is not added.** A strictly-greater sequence gives the ordering; a chain of digests
would give continuity, which nothing here needs yet and which every re-composition would break.
- *Genesis signing its bundle as sequence zero?* Zero is "no order claimed", which is what the bundle
carries by carrying nothing. Same rule, no genesis branch.
- *A marker for a mode change?* Not needed for the incident it guards: a replayed converged declaration
reaching a node returned to adopted is already refused **by mode**, before this check runs.
## How it is checked
Host: an older sequence is refused, a newer or equal one is not, and no order claimed on either side
compares nothing; the drain keeps the highest sequence, and falls back to arrival when none is claimed.
Controller: a send carries its number inside the signed bytes, an unnumbered send is byte for byte what
it was before, and each node's counter is one higher per send and readable for the comparison.
@@ -1,8 +1,8 @@
--- ---
status: located status: resolved
opened: 2026-09-24 opened: 2026-09-24
located-in: [mesh-host internal/apply] located-in: [mesh-controller internal/catalogue/declaration.go (every container was given the roster at creation)]
fixed-by: fixed-by: mesh-controller PR 161 — no container is given a mesh name; it resolves through its machine's resolver (ADR 0148, landed 2026-09-30 once issue 110 did)
amended-design: amended-design:
--- ---
@@ -59,3 +59,27 @@ container is made with are the same kind of input, read once at creation, and ar
and is the stronger statement; it is also what the mesh's own resolver exists for. and is the stronger statement; it is also what the mesh's own resolver exists for.
- Either way: what tells an operator that a container is running with an address the node no longer - Either way: what tells an operator that a container is running with an address the node no longer
has? Nothing did. has? Nothing did.
## Answered at the cause (2026-09-30)
This was the first of three arrivals of one fact: a container is given the mesh's names when it is
created and never looks again, so a name that moves afterwards is wrong inside it for as long as it
runs. It arrived again as [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md),
whose fix made the names comparable — and that fix made the roster part of every container's identity,
which arrived as [issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md).
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) ends the
copying: a container resolves through its machine's resolver at the moment it asks. The shape this
record reports then has nowhere to occur. It is gated on
[issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md), so
until that lands the mesh still copies and still compares.
## Resolved (2026-09-30)
110 landed the same day ([its resolution](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/01-resolution.md)),
and mesh-controller PR 161 then removed the copy: no container is given a mesh name or a mesh address,
and a module's own declared entries are the only `host` lines it carries. Verified on the control-node
after its containers were recreated once — the last time a name will do that: the forge's container
carries no extra hosts and resolves another machine and a routed name through the machine's resolver,
so the shape this record describes has nowhere to occur. Checked in the controller's tests: a
container's declaration is byte-for-byte the same under a roster of one machine and a roster of three.
@@ -1,8 +1,9 @@
--- ---
status: open status: resolved
opened: 2026-09-24 opened: 2026-09-24
located-in: [] located-in:
fixed-by: - mesh-catalog modules/dnsmasq (the runtime was never told; the resolver answered by interface)
fixed-by: mesh-catalog PR 175 (the runtime is reloaded and keeps its containers over a restart) and PR 176 (the resolver answers by address, so a query from a bridge is admitted) — measured 2026-09-30, 01-resolution.md
amended-design: amended-design:
--- ---
@@ -51,3 +52,22 @@ knows that is what the rule means.
the runtime's default one? That is a stronger rule and would have prevented 109 as well. the runtime's default one? That is a stronger rule and would have prevented 109 as well.
- What checks it? A converged bed with a container on the default network resolving a mesh name is - What checks it? A converged bed with a container on the default network resolving a mesh name is
the missing assertion; nothing in the resolver's own beds covers the filter. the missing assertion; nothing in the resolver's own beds covers the filter.
## What now depends on this (2026-09-30)
This stopped being a container-DNS inconvenience.
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) decides
that a container resolves the mesh's names rather than being given a copy of them, which is what stops
one name moving from replacing every container in the mesh
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)) and what makes a
stale address impossible rather than merely noticed
([issues 109](../109-a-container-keeps-the-address-it-was-made-with/00-report.md)
and [135](../135-a-containers-mesh-names-are-not-compared/00-report.md)).
**That decision cannot land until this one does**, and not partly: a container on the runtime's default
network is the case with no DNS at all, and it is the case the mesh's own forge runs in. Two of four
machines also bind the resolver to loopback only, so the runtime hands their containers a public
resolver. Both halves are this issue.
*Later the same day: the second half was wrong, and the first had a different cause than the one above.
[01-resolution.md](01-resolution.md) has what was actually found.*
@@ -0,0 +1,64 @@
# 110 — resolved: a container on any network reaches the resolver, and is answered
*2026-09-30. Measured on the three converged machines; the adopted one holds its resolver module until it
is taken and is not covered.*
## What was actually wrong
Not what the report predicted. The report named the filter: a container on the runtime's default
network asks from a bridge address, and the converged filter admitted queries by source address only.
That was true when it was written and was fixed before this issue was ever tested — the filter admits
by the link a packet arrives on ([ADR 0144](../../02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md)),
and a container's bridge is admitted whole. Tested on every machine: the query arrives, the filter
passes it.
Three other things were wrong, each hiding the next.
**The runtime had never been told.** The resolver module writes the runtime's `dns` key into the
runtime's own configuration file. The runtime reads that key when it starts and not on a reload, and on
two machines the runtime predated the file — so every container they started got a public resolver, and
`novox.internal` came back as not existing. Nothing reported this: the file was present and current,
the resolver ran, and a name not existing is a valid answer. Fixed in mesh-catalog PR 175: the module
also sets `live-restore` and reloads the runtime when its file changes, so the one restart the `dns` key
needs no longer stops every container. The restart is then the operator's, once per machine; done on
both today, with every running container kept.
**The resolver dropped the query.** With the runtime corrected, a container's query reached the resolver
— and got no answer, on every machine, including the one whose runtime had been right all along. The
socket was bound to the private address; the filter admitted the packet; dnsmasq received it and
discarded it without a line of log. Its configuration said `interface=mesh0`, and dnsmasq admits a
query by the interface it arrives on when told an interface: a container's query is addressed to the
private address but arrives on the runtime's bridge, and the bridge is not `mesh0`. Fixed in mesh-catalog
PR 176: the resolver is told the address to answer on, not the interface that carries it, and a query to
that address is admitted whatever bridge brings it. The bridges are the runtime's to name.
**The report's second half was wrong.** "Two of four machines bind the resolver to loopback only" was
an inference from the containers' behaviour, and the behaviour had the cause above. The resolver bound
the private address on all four; nothing had asked it there.
## What is verified
From a container on the runtime's default network, started by hand and given nothing, on each of the
three converged machines: `novox.internal` answers with the hub's private address, through the machine's
own resolver. That is the fourth check of
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) — "on every
network the runtime offers" — and its first step; the record's step 2 (the runtime told per machine, as
a file) was already how the module works. Step 3 may now begin.
## What checks it
By hand, today. Nothing in the mesh asserts that a container can resolve a mesh name: the resolver's
own tests cover what it answers, not who can ask. The check that would have caught all three faults is
the one the report asked for and 0148 lists — a container on the default network resolving a mesh name
— and it is not built. It belongs with the reachability check of
[issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
which is parked; until then this is a thing a person verifies after touching the resolver, the filter,
or the runtime's configuration.
## What this cost to find
The three faults produced one symptom — a container that cannot resolve — and each fix revealed the
next. The first was found by reading the runtime's own view of its configuration rather than the file;
the second by capturing the query on the bridge and finding it arrive and go unanswered; the third only
by admitting the first belief was wrong. A machine that had been believed to work all day had never
worked either.
@@ -1,5 +1,5 @@
--- ---
status: fixed status: resolved
opened: 2026-09-24 opened: 2026-09-24
located-in: [mesh-controller internal/inventory, mesh-controller internal/catalogue, mesh-controller cmd/mesh-controller, mesh-catalog modules/dnsmasq] located-in: [mesh-controller internal/inventory, mesh-controller internal/catalogue, mesh-controller cmd/mesh-controller, mesh-catalog modules/dnsmasq]
fixed-by: [mesh-controller#73 overlay-name + namesInTheMesh, mesh-catalog#106 dnsmasq daemon.json merge] fixed-by: [mesh-controller#73 overlay-name + namesInTheMesh, mesh-catalog#106 dnsmasq daemon.json merge]
@@ -1,8 +1,8 @@
--- ---
status: open status: resolved
opened: 2026-09-24 opened: 2026-09-24
located-in: [mesh-controller module.json, mesh-host internal/apply] located-in: [mesh-controller module.json, mesh-host internal/apply]
fixed-by: fixed-by: ADR 0142 — the mesh's own components are delivered as binaries on the machine, so the controller is a process; the delivery itself is issue 142
amended-design: amended-design:
--- ---
@@ -84,3 +84,21 @@ restart and run-to-completion semantics — so this would not need host-side wor
The two do not collapse into one. The controller is not a code-carrying sidecar, and `network: host` The two do not collapse into one. The controller is not a code-carrying sidecar, and `network: host`
is what makes the asymmetry visible here and nowhere else. is what makes the asymmetry visible here and nowhere else.
## Answered
*2026-09-29, in a grooming pass.* This asked a question rather than reporting a defect, and the
question was taken: [ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md)
decides that the mesh's own components — the host, the controller, the catalogue, the builder, the
vault — are **binaries on the machine**, delivered by the mechanism
[ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) built, and that
third-party software (the store, the registry, the broker) stays a container because an image is the
right way to carry somebody else's build.
So the operating experience this record was written from — every mutating command reached through
`docker exec mesh-controller` — is answered, and answered against the container.
**The delivery is a separate matter and is not this record's.** Step 1 of it is built and no
component travels yet; that is
[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md).
@@ -1,8 +1,8 @@
--- ---
status: located status: resolved
opened: 2026-09-25 opened: 2026-09-25
located-in: [hq, mesh-catalog modules/showcase, mesh-sdk src/tools/index.ts, mesh-tools] located-in: [hq, mesh-catalog modules/showcase, mesh-sdk src/tools/index.ts, mesh-tools]
fixed-by: fixed-by: hq 83791f0 (PR 196) — ADR 0150: a module's own code runs as supervised processes under the module's one account; designs 18 and 20 now cite it, and ADR 0047 carries a dated note pointing at it
amended-design: amended-design:
--- ---
@@ -99,3 +99,27 @@ what the sidecar is has nowhere in the design layer to look, which is how this w
is mechanically checkable: the resource types a design doc names are a closed set, and every is mechanically checkable: the resource types a design doc names are a closed set, and every
member of it either appears in a decision or does not. Whether that check is worth writing is member of it either appears in a decision or does not. Whether that check is worth writing is
part of this issue, not settled by it. part of this issue, not settled by it.
## Answered (2026-09-30)
[ADR 0150](../../02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md)
settles all three disagreements, and the design documents win two of them:
1. **Container or unit — a supervised process.** ADR 0047's argument never required a container. It
argued for a runtime *per module*, because a node-wide one could not hold a per-module account and
per-module runtimes on one tool key would be handed calls for tools they do not have. A unit per
module satisfies that exactly, and a unit runs as an account. The container was the mechanism to
hand. [ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md) had
already gone the same way for the mesh's own components, and a module is not a container.
2. **One process or several — several, under one account.** 0047's "one module, one process, one
account" carried its weight in the last clause; its stated worry was "not a second one to scope and
seal", which is about a second *identity*. Processes sharing the module's one account create none.
What a module may not have is two accounts.
3. **Whether the record was consulted — fixed rather than answered.** Designs 18 and 20 now name 0150
in `decisions:`, and 0047 carries a dated note saying where its hosting form was settled, so neither
door leads to the wrong answer any more.
**What this does not fix.** A module's code delivered as a binary is behind the same gap as the host's
own ([ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md), accepted and not
built): a container's code arrives by `docker pull` and this does not, so until delivery exists such a
module is one somebody places by hand. 0150 records that as the cost of the decision, unpaid.
@@ -1,8 +1,8 @@
--- ---
status: located status: resolved
opened: 2026-09-25 opened: 2026-09-25
located-in: [mesh-catalog modules/umami] located-in: [mesh-host internal/apply (the spec comparison, via issue 135) — not mesh-catalog modules/umami, which this record first named and which was never at fault]
fixed-by: fixed-by: mesh-host e82789a (issue 135's fix) — recreating the container with a current roster ended it; verified live 2026-09-30
amended-design: amended-design:
--- ---
@@ -0,0 +1,79 @@
# 118 — resolved: it was issue 135, and it is over
*Verified on the machine, 2026-09-30.*
## It no longer happens
```
$ docker inspect -f '{{.RestartCount}}' umami
0
$ docker logs umami --tail 12
26 migrations found in prisma/migrations
No pending migrations to apply.
✓ Database is up to date.
▲ Next.js 16.3.4
✓ Ready in 0ms
$ curl -o /dev/null -w '%{http_code}' https://umami.novox.be/
200
```
No restart loop, no `prisma.$queryRaw()` timeout, and the public name that had answered `502` since
2026-09-25 answers `200`. The raw query that could not complete now runs twenty-six migrations and
reports the store up to date.
## What it was
**The same fault as [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md), which
was diagnosed three days later without either record noticing the other.** 135's container is this
one, named in its own evidence table:
```
umami created 2026-09-23 novox.internal:10.42.0.1
mesh-catalog created today novox.internal:10.10.0.1
```
The mesh's overlay range had moved. Umami had been created before the move and held the store's name
at an address that no longer existed, while every container made after the move held the current one.
That is why the dial appeared to succeed and the first real query timed out, and why the same query
from the same network with the same credential answered in milliseconds — **what differed was the
name, not the path, not the credential and not the store.**
Today it holds `novox.internal:10.10.0.1`.
## Why this record did not find it
The report's own reasoning is worth keeping as a warning, because it is careful and it is wrong:
> A dial that succeeds and a query that times out, from a container on one network to a store on
> another, has the shape of a path-MTU or conntrack fault (large response packets dropped after the
> small handshake ones pass), or of the store accepting the TCP connection while the backend it
> proxies for is wedged.
Both are good hypotheses about a network path. Neither is the answer, and the record also names the
move that would have found it — *"what is known to differ for umami against every working consumer of
the same store tonight is nothing yet — that comparison is the first move"* — and then did not make it.
Comparing umami's hosts entries against any container created that week would have shown a five-day-old
address in one field.
**A stale name presents as a network fault.** That is the lesson, and it is the reason
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) stops
copying names into containers at all: not because detecting staleness is hard, but because it disguises
itself as something else for five days while every check reports success.
## What actually ended it
Issue 135's fix — `mesh-host` e82789a, *a container's mesh names are part of what it is* — put the
roster into the spec digest the host compares, so a container whose names moved is recreated like one
whose image moved. That recreated umami with a current roster and ended this.
That fix has since been superseded in turn, by 0148, because making the roster part of every
container's identity meant one name moving replaced every container in the mesh
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). So this record
closes on a fix that is itself on the way out — which does not make it less closed, and is worth saying
plainly rather than leaving a reader to find out.
## Not carried forward
The `502` had one other contributor worth recording as ruled out: a stale duplicate Traefik router for
this name, hand-authored in the predecessor's era beside the mesh-written one, removed on 2026-09-25
and — as the report says — changing nothing. It was not the cause and it is gone.
@@ -1,8 +1,8 @@
--- ---
status: located status: resolved
opened: 2026-09-26 opened: 2026-09-26
located-in: [mesh-sdk src/provisioner, mesh-catalog modules/redis] located-in: [mesh-sdk src/provisioner, mesh-catalog modules/redis]
fixed-by: fixed-by: mesh-catalog bbda88c, merged in #84 — every credential provider says whether it still holds a consumer
amended-design: amended-design:
--- ---
@@ -60,3 +60,9 @@ checks it after the first pass.
instance and leaves the gap for the others. instance and leaves the gap for the others.
- Where does the record of what was applied live, if not in memory? ADR 0114, still - Where does the record of what was applied live, if not in memory? ADR 0114, still
proposed, puts rotation state with the vault. The same place may answer this. proposed, puts rotation state with the vault. The same place may answer this.
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A provisioner asks the backend what is there rather than trusting what it remembers doing. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -1,7 +1,9 @@
--- ---
status: open status: resolved
opened: 2026-09-26 opened: 2026-09-26
located-in: [mesh-host internal/apply, mesh-controller] located-in: [mesh-host internal/link (the apply report line), mesh-controller cmd/mesh-controller (status and its all-well condition)]
fixed-by: mesh-host cbf5018, mesh-controller bfd983e
amended-design:
--- ---
# 125 — a hold is not a line in the apply report, and an operator flew blind into an outage # 125 — a hold is not a line in the apply report, and an operator flew blind into an outage
@@ -0,0 +1,70 @@
# 125 — resolved: a hold is a line in the report, and it stops the mesh reading as well
*2026-09-30.*
## What the four surfaces say now
The report named four surfaces, none of which carried the one sentence that mattered. Two of them
already did by the time this was picked up, and two did not.
**1. The apply report says what it held** — this was missing, and is the line the operator was reading
when the count did not add up:
```
applied 330 resource(s), 16 held until their module is taken (route-proxy: 13, mailu: 3)
```
Grouped by module and ordered by name, because `take` acts on a module and that is the sentence an
operator needs. An apply that held nothing says nothing extra — a line reporting `0 held` on every
converged apply is one that stops being read.
**2. `status` counts holds, and a hold breaks "all well"** — this was missing. Status now says:
```
15 resource(s) are held as found, because their module was assigned and never taken — so it is
running none of what it declares:
novox route-proxy (13), mailu (2)
`take <node> <module>` compares what runs against what it declares, and runs it
```
**And the sentence that was the fault no longer prints.** "all doing what they were told, all heard
from, running what the mesh would send them" was *true* for the whole outage, and acting on it stopped
the predecessor's proxy. A hold now suppresses it; being adopted still does not, and the difference is
deliberate — adopted is a mode somebody chose and can leave alone, a module assigned and never taken is
a half-finished action with nothing left to finish it.
**3. `node show <node>` shows the node's own held list** — already true, and recorded here as checked
rather than assumed. It reads `held` from what the machine last reported, with an `as of` beside it, and
names each hold's kind, target, id, module, whether something other than the mesh has changed it, and
where an original was kept.
**4. The push's count** is unchanged and now interpretable, which was the ask: `sent 346` against
`applied 330, 16 held until their module is taken (…)` is a pair a reader can resolve without opening a
file on the machine.
## Where the data came from
**The host already reported it.** `Held` has been on the wire since [ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md),
and the controller already recorded it and showed it in `node show`. Nothing needed a new field, a new
message or a migration — which is why this is additive, and why the report's framing (*"the semantics
are consistent and right; the reporting is what let them be forgotten"*) was exactly right.
What was missing was that two surfaces never asked. Status read the mesh's take-time listing, so a
module assigned after that listing showed nothing at all; the apply line counted what it applied and
said nothing about the difference.
## One thing deliberately not done
**Status does not call a hold a fault.** It is correct behaviour, and a reader trained to see red for
something the mesh did right will stop reading. It is reported as work outstanding, with the command
that finishes it — and it withholds the all-well sentence, which is the part that carries the weight.
## How it is checked
- The apply line names the count and the module, and is empty when nothing is held (mesh-host).
- Status finds a hold from what the machine reported, end to end through the store.
- **A held module makes the all-well condition false**, asserted against the production condition
rather than a copy of it — that condition is now one named function for this reason.
- The JSON form carries a row per machine and module, and omits the field entirely when nothing is
held.
@@ -0,0 +1,93 @@
---
status: resolved
opened: 2026-09-27
located-in: [mesh-catalog modules, mesh-control internal/catalogue, mesh-tools src]
fixed-by: mesh-catalog 7b06a7a, mesh-tools fbeb373, mesh-control 05ff606
amended-design: 03-DESIGN/01-to-be/32-what-a-module-declares.md
---
# 127 — A module's event derives a subject nothing publishes
## What was observed
[Design 29](../../03-DESIGN/01-to-be/32-what-a-module-declares.md) §1 says a module names an event
locally and the mesh derives the subject: `emits: order.placed` becomes
`mesh.mod.<module>.event.order.placed`, and a consumer declaring `consumes: shop.order.placed`
subscribes the emitter's own subject. That derivation is built and tested.
**Every event name in the catalogue is still written the way a routing key on the bus the mesh has
is written** — `module.<module>.<verb>` — and the derivation reads it as `<emitter>.<event>`. Asked
of the composer directly, with the module names and declarations the catalogue holds today:
| declared | derived |
|---|---|
| `builder` emits `module.builder.built` | publish `mesh.mod.builder.event.module.builder.built` |
| the catalogue consumes `module.builder.built` | subscribe `mesh.mod.module.event.builder.built` |
| a media module emits `module.<itself>.download.completed` | publish `mesh.mod.<itself>.event.module.<itself>.download.completed` |
| a player consumes `module.*.download.completed` | subscribe `mesh.mod.module.event.*.download.completed` |
The consumer's subject names a module called `module`. **No cross-module subscription in the
catalogue matches what any emitter publishes.** Thirty-seven manifests declare events; every one of
their consume declarations derives this way.
Two further consequences of the same cause, found in the same check:
- One module declares `consumes: "#"` — the wildcard of the bus the mesh has, which is not a
subject at all. The composer **refuses it outright**, so that module's account cannot be composed
and the module cannot be assigned.
- One module emits under a name that is not its own — it declares `module.<other>.image.pushed`
while being a differently named module — which the derivation puts inside *its* namespace. Whether
that is legitimate is a design question: design 32 §2 makes an event's source a fact the server
enforces, and this is a module claiming another's name in its own event.
None of it fails on the bus the mesh runs on today, where a routing key is matched literally and
nothing derives anything. It fails only once the subject is derived — which is to say it fails on
the first mesh raised on the new bus, and not before.
Evidence: run against the controller's own `PermissionsFor` on the current feature branch, with the
declarations read from the catalogue's manifests. Found while wiring the controller's consume side
(design 28 step 3.4), when the controller's own subscription had to be written and the subject it
would have to name turned out not to be the one design 32 specifies.
## Why it matters beyond this instance
**This is the failure [ADR 0074](../../02-DECISIONS/0074-the-wire-is-specified-not-the-types.md)
exists to catch, arriving by a route the conformance suite does not cover.** Two implementations
that disagree about an envelope do not fail to compile — they ignore each other while both keep
running. Here it is not two implementations disagreeing but a *declaration* and a *derivation*
disagreeing, and the symptom is identical: every service starts, every log is quiet, and nothing
reacts to anything.
The fixtures cannot catch it. They pin one emitter's envelope against one subject, and both halves
of that pair are correct. What is wrong is only visible when an emitter's derived subject is set
beside a consumer's derived subject — a check nothing performs, because until the subject was
derived there was nothing to compare.
It also means the rule 3.8 established is weaker than it reads. That task asserted **no manifest
contains a subject**, which holds: a manifest contains a local name. What nothing asserts is that a
local name derives to a subject some emitter actually publishes, and the rule as stated is satisfied
by thirty-seven manifests whose names derive to nothing.
And it blocks work already scheduled. Design 28's step 4.2 (a build source's change reaching the
builder over the bus) and 4.3 (an installation completing over the bus) are both event flows through
exactly these pairs, and the catch-up flow the controller answers is a third — the controller
currently replays a build announcement under its *own* name rather than the builder's, which a
consumer filtering the builder's subject will not hear either.
## Open questions
- Is a local name converted per manifest (`emits: built`), or does the derivation keep accepting the
old form and strip a redundant prefix? The first is thirty-seven manifests and a rule that can be
checked; the second is a rule that cannot, because `module.foo.bar` is also a legitimate three-part
local name.
- What checks the pair? An emitter's derived subject against every consumer's derived subject is a
whole-catalogue check, not a per-manifest one — and a module lives in its own repository and may
be registered long after the catalogue was checked.
- What are `#` and `*` in a consumed name? The bus the mesh has and the bus being built spell
wildcards differently, and a `consumes` pattern is the one place a module writes one.
- May a module emit an event named after another module, and if not, what does the module that does
it today declare instead?
- Who replays? A catch-up answer published by the controller under a builder's subject is the
controller signing an event as another module, which is the thing the derived namespace prevents.
If it must not, then a replay is a different message from an announcement, and the consumer needs
to be told so.
@@ -0,0 +1,109 @@
# Diagnosis — 2026-09-27
## Where it lives
Three places, and only one of them is a bug in code.
**The manifests, in the module catalogue.** Thirty-seven declare events, and every one of them
spells an event the way a routing key on the bus the mesh runs on today is spelled —
`module.<module>.<verb>`. [Design 29](../../03-DESIGN/01-to-be/32-what-a-module-declares.md) §1 says
a module names an event **locally and bare** (`emits: order.placed`) and a consumer names
`<emitter>.<event>` (`consumes: billing.order.placed`). So the manifests are stale against a rule
that was already decided, not wrong against an undecided one. **This is the whole of the reported
symptom.**
**The manifest's own documentation, in the parser.** The comments on `emits` and `consumes` still
describe the old convention and give the old examples — "dotted topic keys, e.g.
`module.umami.site.created`", and `"#"` named as the audit logger's pattern. A module author reading
the file they read most is being told to write the thing that does not work. That is why the drift
was uniform across thirty-seven manifests rather than scattered: nobody was mistaken, everyone
followed the documentation.
**Nothing checks either one.** `ParseManifest` validates the module name, the slug, what it provides
and what it requires. It says nothing about an event name. So a local name that derives to a
namespace belonging to a module called `module` is accepted by every check the mesh has, and the
first thing that notices is a subscription that never fires.
## What was ruled out
**The derivation is not wrong.** Asked directly, with the module names and declarations the
catalogue holds, `PermissionsFor` produces exactly what design 32 §1 specifies for the input it is
given: it reads a consumer's `<emitter>.<event>` and builds the emitter's subject. Given
`module.builder.built` it reads the emitter as `module`, which is a correct reading of an incorrect
declaration.
**The conformance fixtures are not at fault and could not have caught it.** They pin one emitter's
envelope against one subject, and both halves of that pair are correct. What is wrong is only
visible when an emitter's derived subject is set beside a *consumer's* derived subject — a
comparison nothing performs, because until the subject was derived there was nothing to compare.
**Task 3.8's rule is not broken, it is weaker than it reads.** That task asserted **no manifest
contains a subject**, which holds: a manifest contains a local name. Nothing asserts that a local
name derives to a subject some emitter actually publishes.
## What is still a decision and not a conversion
Converting the manifests is implementing design 29, not deciding anything. Three of the report's
open questions are not:
- **Wildcards.** Design 29's table has no wildcard row, and two manifests need one: a module
consuming every download completion across several media modules, and an audit logger consuming
everything. The two buses spell wildcards differently, and a `consumes` pattern is the one place
a module writes one.
- **A module emitting under another module's name.** One manifest declares an event named for a
*provision* rather than for itself. Design 29 §2 makes an event's source a fact the bus enforces,
so this cannot survive as written — and the remedy is probably not a rename but a **seat**, which
is what a name stable across whoever implements it already is.
- **Who replays a build announcement.** The controller answers a catalogue's catch-up by
re-publishing builds under its *own* name, which no consumer of the builder's subject hears, and
for which it holds no grant. Publishing them under the builder's subject would be the controller
signing an event as another module — the exact thing the derived namespace prevents. So the
catch-up is either a different message or a different mechanism, and that is a decision.
## Owners
`located-in` names the manifests and the parser. The replay question reaches the controller and the
catalogue module together and is recorded above rather than in that field, because it is not where
this symptom lives.
# Fixed — 2026-09-27
Converted, and the rule now has checks. What it took was larger than the report said, in two
directions nobody had looked.
**The module code, not just the manifests.** Forty-three files pass an event name to `emit()` at
runtime, and the runtime builds the subject from what it is handed. A converted manifest with
unconverted code would have had the permission and the subject disagree — the same silence, one layer
down.
**Both clients had to learn the mapping.** Each passed the name straight through, which was right on
the bus the mesh runs on today only because modules were writing routing keys. So the old bus's client
now turns a local name into `module.<emitter>.<event>` on the way out and back on the way in.
**Without that, converting the modules would have broken the mesh that is actually running** — which
is the opposite of what fixing this was for.
**The declaration and the handler spoke different vocabularies.** The key a module's handler saw was
the event name alone, while its manifest names `<emitter>.<event>`. So a correct manifest produced a
pattern that could never match. The subject already carries the emitter; the key names it now.
## The three open questions, answered
- **Wildcards**: `*` is one name, `**` is the rest, spelled the mesh's way and derived to each bus's
own. `**` alone is every event, which is what the audit logger wanted and now says in one token.
- **A module emitting under another's name**: not allowed, and the remedy is the seat rather than a
rename — a role's name outlives whoever fills it. **Deferred in practice**: seats carry protocol in
the manifest and in the permission model, and the shared library cannot publish on one, so the
module that did this emits under its own name and its consumers carry that coupling. Worth a task
when a seat's holder needs to emit.
- **Who replays a build announcement**: still open, and narrowed. It cannot become a reply to the
catalogue's inbox: answering a module's inbox needs `_INBOX.>`, which is the blanket grant
[design 25](../../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §4 refuses. So the remaining options are
a subject the controller may publish and the catalogue may subscribe, or a durable consumer that
starts at the beginning of the stream and removes the need to ask at all. Recorded on the work
breakdown as the catch-up half of task 4.5 rather than here, because it is no longer this symptom.
## What it found while running
Two dangling subscriptions that predated this and nothing had reported: a module emitting an event its
manifest never declared, which the new bus refuses outright, and a module waiting for an event nothing
emits — a demo that could never be triggered, because only that module may publish under its own name.
@@ -1,5 +1,6 @@
--- ---
status: located status: resolved
fixed-by: mesh-host 1cb8953 and fdc768c — the mesh writes into a marked block of a text file instead of over it
opened: 2026-09-26 opened: 2026-09-26
located-in: [mesh-controller internal/catalogue/facts.go, mesh-host internal/apply] located-in: [mesh-controller internal/catalogue/facts.go, mesh-host internal/apply]
--- ---
@@ -71,3 +72,9 @@ private network loses that name too.
- The host's file resource supports `into: "json"` only; anything else is a whole write. - The host's file resource supports `into: "json"` only; anything else is a whole write.
- `node show <node>` on the adopted workstation: `holds file /etc/hosts - `node show <node>` on the adopted workstation: `holds file /etc/hosts
mesh-wireguard.fact-node-names`, original kept. mesh-wireguard.fact-node-names`, original kept.
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A shared hosts file keeps every line that is not the mesh's. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -1,7 +1,8 @@
--- ---
status: open status: resolved
opened: 2026-09-26 opened: 2026-09-26
located-in: [mesh-controller, mesh-catalog step-ca] located-in: [mesh-catalog modules/ca-trust]
fixed-by: the ca-trust module registered from the catalogue and assigned to a workstation; verified in both directions 2026-09-30 (02-resolution.md)
--- ---
# 129 — nothing makes a machine trust the mesh's own certificate authority # 129 — nothing makes a machine trust the mesh's own certificate authority
@@ -0,0 +1,82 @@
# 129 — diagnosis
*2026-09-30, from the workstation the issue was opened on.*
## Still live, and reproduced exactly
The certificate is genuine, the authority is the mesh's, and nothing on the machine trusts it:
```
$ openssl s_client -connect keycloak.novox.internal:443 -servername keycloak.novox.internal
subject=CN=keycloak.novox.internal
issuer=O=Mesh Internal CA, CN=Mesh Internal CA Intermediate CA
Verify return code: 20 (unable to get local issuer certificate)
$ curl https://keycloak.novox.internal/
curl: (60) SSL certificate OpenSSL verify result: unable to get local issuer certificate (20)
```
`trust list` holds no entry for the mesh. The anchors present are two `mkcert` development roots and
the predecessor's lab root — the report's account of the trust store is unchanged.
The public name on the same proxy verifies cleanly (`CN=keycloak.novox.be`, Let's Encrypt, return code
0), which places the fault exactly where the report puts it: not in the proxy, not in the authority,
and not in the certificate.
**The name matters, and the report's "every HTTPS name the mesh serves internally" is too broad.** The
served internal name is `<label>.<node>.internal`. The hosts file also carries
`<label>.<public-domain>.internal`, which nothing serves and which fails differently — that is
[issue 157](../157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md), found while
reproducing this, and it cost the first several minutes of this diagnosis.
## The authority serves what the module needs
`step-ca` is up and healthy, and publishes `roots: /roots.pem` for both `acme-ca` and
`internal-acme-ca`. That endpoint returns PEM:
```
$ curl -sk https://127.0.0.1:9000/roots.pem
-----BEGIN CERTIFICATE-----
MIIBvzCCAWWgAwIBAgIQYa2CkdJk16JyG/dVy2qoEzAKBggqhkjOPQQDAjA+…
```
So `${bound:internal-acme-ca:roots}` in the `ca-trust` module composes to a URL that returns a
certificate, and the module's own check — refuse a body that is not one — is checking the right thing.
Worth recording because it was nearly filed as a defect: step-ca *also* serves `/roots`, which returns
`{"crts":["-----BEGIN CERTIFICATE-----\n…"]}`. That body contains the literal text the module greps
for, so had the module used `/roots` it would have installed JSON into the anchors directory and
reported success. It does not use it. The guard is sound only because the published path is the PEM
one, which is worth knowing before anybody changes either.
## What is actually in the way
**The module is not registered.** The report and the work plan both say it exists and is merged, which
it does — `mesh-catalog modules/ca-trust`, on `main`. But the mesh has never been told about it:
```
$ mesh-controller module list | grep -iE 'ca-trust|step-ca'
step-ca 1 built 67f5f4cf on novox
```
39 of the catalogue's 76 manifests are registered. `ca-trust` is one of the 37 that are not, so it
cannot be assigned to anything — "assign it to one machine" has no module to name.
A dry run confirms it registers cleanly and needs no artifact built: it declares a directory, a script,
a unit and a service, and no image.
```
$ mesh-controller build <catalogue> --path modules/ca-trust --dry-run
… the manifest, parsed and validated
```
## So the remaining work is three steps, not one
1. **Register it** — build it from the catalogue, which pins nothing because it has no artifacts.
2. **Assign it** to a machine. The workstation this was observed on is the honest first choice: it is
where a person meets the fault, and it is where the check can be made with a plain client.
3. **Verify** `curl https://<label>.<node>.internal/` with no flags, and `trust list` naming the mesh.
Then removal, which the module declares and nothing has exercised: unassigning must take the anchor
away and refresh the bundles ([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md)),
and that is the half most likely to be wrong, because it is the half nobody reaches by accident.
@@ -0,0 +1,104 @@
# 129 — resolved: the workstation trusts the mesh, and stops when told to
*Done and measured on the machine, 2026-09-30.*
## What was done
Three steps, not the one the plan expected — the module was merged and had never been registered
(see [the diagnosis](01-diagnosis.md)):
1. **Registered** `ca-trust` from the catalogue. No artifact to build: it declares a directory, a
script, a unit and a service, and no image.
2. **Assigned** it to the workstation the issue was opened on, and pushed.
3. **Verified** with a plain client, then **unassigned and pushed again** to exercise removal, then
assigned and pushed once more.
The host's own account of arriving:
```
created ca-trust.state (/var/lib/ca-trust)
created ca-trust.anchor (/var/lib/ca-trust/anchor)
created ca-trust.unit (/etc/systemd/system/mesh-ca-trust.service)
updated ca-trust.trust (mesh-ca-trust.service): boot disabled to enabled, stopped to running
```
## It works, by the check the report asked for
The report's own reproduction, with no flags and nothing installed by hand:
```
$ curl -sS -o /dev/null -w '%{http_code}' https://git.novox.internal/
200
$ openssl s_client -connect git.novox.internal:443 -servername git.novox.internal
issuer=O=Mesh Internal CA, CN=Mesh Internal CA Intermediate CA
Verify return code: 0 (ok)
$ trust list | grep -A2 Mesh
label: Mesh Internal CA Root CA
trust: anchor
category: authority
```
Four internal names, all verifying: `git` 200, `umami` 200, `keycloak` 302, `drive` 302. Before this,
every one of them was `curl: (60) … unable to get local issuer certificate (20)`.
**And the consequence the report named specifically**: git over HTTPS to the mesh's forge, which it said
had forced the working clone URL to be ssh-only.
```
$ git ls-remote https://git.novox.internal/novox/hq.git HEAD
76fbe323ea3401fcdadbf500c61bd3fa5a0a8603 HEAD
```
## Removal is symmetric, which nothing had ever shown
[ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md) says the module anchors the
authority *and takes it away again*. That half had never run. Unassigning and pushing:
```
removed ca-trust.trust (mesh-ca-trust.service)
removed ca-trust.unit (/etc/systemd/system/mesh-ca-trust.service)
removed ca-trust.anchor (/var/lib/ca-trust/anchor)
removed ca-trust.state (/var/lib/ca-trust)
```
Then: the anchor file gone, `trust list` naming no authority of the mesh's, and the plain client back to
`unable to get local issuer certificate (20)`. A machine that leaves the mesh stops trusting it, as the
record claims.
**The order is what makes it work, and is worth saying.** The service is removed *first*, so systemd
runs the unit's `ExecStop` — which is what deletes the certificate and refreshes the bundles — while the
script it calls still exists. Had the script or the state directory gone first, stopping the unit would
have had nothing to run, and the anchor would have been left behind with nothing declaring it. Nothing
in the module says this; it is the host's removal order that makes the module's symmetry real.
## What this leaves
- **Three machines of four** *(extended the same day, after the first was proven)*. `ca-trust` is now
assigned to every converged machine, and each verifies with a plain client:
```
novox mesh CA in trust store: 1 https://git.<node>.internal/ -> 200
g14 mesh CA in trust store: 1 https://git.<node>.internal/ -> 200
shanks mesh CA in trust store: 1 https://git.<node>.internal/ -> 200
```
Before, each of the two that had not been assigned it answered
`curl: (60) … unable to get local issuer certificate (20)` and held no entry for the mesh.
**`ace` is deliberately not among them.** It is the adopted machine, still carrying the
predecessor's resolver and filter, and it is not converged until the network and module-assignment
work is settled. Assigning a module to it would hold rather than run
([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)), which is
correct and is not the same as trusting anything.
- **It arrives per assignment, which is a shape worth questioning.** The report's reasoning — that
being on the private network is what makes a machine one that speaks to the mesh's names — argues
for a fact carried to every machine on the network, the way the roster and the registry trust are.
ADR 0147 chose a module assigned per machine, and this record does not reopen it; the cost is that
a machine joining the mesh trusts nothing until somebody remembers a second command.
- **`service-manager` reports `degraded` on this machine** and the module ran anyway. Worth knowing that
the capability gate passes on a degraded service manager, since a module whose whole delivery is a
unit is the kind that would be worst served by one.
- **The predecessor's authority is still in the trust store**, beside the mesh's now rather than instead
of it. The report notes that it cannot be retired while anything on the machine speaks TLS to a mesh
name; that is no longer true here, and retiring it is its own piece of work.
@@ -1,5 +1,6 @@
--- ---
status: located status: resolved
fixed-by: mesh-host 3112c88 — undeclaring gives a unit back the state it was found in, and removes only a process the mesh made
opened: 2026-09-27 opened: 2026-09-27
located-in: [mesh-host internal/apply/apply.go (remove)] located-in: [mesh-host internal/apply/apply.go (remove)]
amended-design: 02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md amended-design: 02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md
@@ -9,13 +10,13 @@ amended-design: 02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was
## What was observed ## What was observed
Reviewing the uplink modules ([ADR 0117](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md)) Reviewing the uplink modules ([ADR 0125](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md))
found that the host's `remove` path stops every `service` resource that is no longer declared: found that the host's `remove` path stops every `service` resource that is no longer declared:
`SetServiceState(..., "stopped")`, reported as "stopped; the unit file is not the host's to `SetServiceState(..., "stopped")`, reported as "stopped; the unit file is not the host's to
delete". `store.Orphans` matches by id alone. So any of these stops the unit: delete". `store.Orphans` matches by id alone. So any of these stops the unit:
- the module is unassigned — by mistake, or to switch it for another; - the module is unassigned — by mistake, or to switch it for another;
- the node is sent a deliberately-empty declaration ([issue 127](../127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)); - the node is sent a deliberately-empty declaration ([issue 149](../149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md));
- a later catalogue version renames the resource's `id`. - a later catalogue version renames the resource's `id`.
That is right for a service the mesh brought into being. It is wrong for a unit the mesh That is right for a service the mesh brought into being. It is wrong for a unit the mesh
@@ -45,7 +46,7 @@ reaches it by.
## Resolution ## Resolution
[ADR 0118](../../02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md): [ADR 0126](../../02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md):
undeclaring removes what the mesh made and gives back what it changed. The host records the state undeclaring removes what the mesh made and gives back what it changed. The host records the state
it first found a unit in, and undeclaring returns the unit to it — a unit found running (the it first found a unit in, and undeclaring returns the unit to it — a unit found running (the
container runtime, sshd, a network manager) is left running; one the mesh started (the packet container runtime, sshd, a network manager) is left running; one the mesh started (the packet
@@ -64,3 +65,9 @@ something to settle in passing.
The unassign preview is partly answered — the host's plan names each unit it will stop — and the The unassign preview is partly answered — the host's plan names each unit it will stop — and the
controller's side is left open. controller's side is left open.
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* Undeclaring no longer stops a unit the mesh only reloaded or only kept running. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -0,0 +1,107 @@
---
status: resolved
opened: 2026-09-27
located-in: [mesh-catalog modules/gitea, mesh-controller cmd/mesh-controller, mesh-controller internal/builder]
fixed-by: mesh-catalog #124 — the forge module watches for merged pull requests and emits `pull.merged` with the merge commit and the clone address; mesh-controller #110/#111 — the control plane follows that event on the bus, marks every module built from that repository and branch as moved, and builds them bases first, stopping when a base fails; mesh-controller #113/#114 — a build records the bases it was handed and the graph is read from builds, without which "bases first" had no edges to order by.
amended-design: 03-DESIGN/01-to-be/28-building-the-bus.md
---
# 131 — Nothing tells the mesh a source moved, and it reports itself current anyway
## What was observed
Six changes were merged to the trunk of six repositories in one sitting. The build machine built
nothing. Its last build, minutes before the first merge, was still the one it reported; no build was
requested, refused or failed, because none was ever asked for.
Asked afterwards what was wrong, the mesh said:
> 4 machine(s), all doing what they were told, all heard from, running what the mesh would send them,
> and every module current with its source
Every one of the six had moved. The last clause was false for all of them, and it is the clause a
person reads to decide whether there is anything to do.
## Why it matters beyond this instance
**The mesh learns a source moved by being told, and there is no longer anything to tell it.** The
command exists — a person names the module and the commit — and so does the question the overview
answers. What is missing is whatever used to connect the two. One repository still carries a forge
webhook aimed at a port; the rest carry none, and the port belongs to a different service than the
one the arrangement implies. So the state is not "the trigger is broken" but "there is no trigger,
and nothing says so".
**A wrong answer is worse here than no answer.** "Every module current with its source" is
indistinguishable, to a reader, from a mesh that has genuinely caught up. The overview is built to be
the thing you check instead of checking by hand, so a confident false negative removes the habit that
would otherwise have caught it. Nothing in the mesh is at fault for being out of date — it is at
fault for saying it is not.
**It is also why "current with its source" cannot be a stored fact.** The mesh compares what it built
against what it was last told the source was, and calls that agreement. Two facts agreeing tells you
nothing when both come from the same place.
## The intended shape, which is decided and not built
The forge emits what happened to it — a pull request merged — and the build machine reacts by
building what that commit affects. That keeps the forge ignorant of the build system and the build
machine ignorant of the forge's internals, which is the same argument
[ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) makes for addressing an event
to its emitter: the merge is a fact about the forge, and what should be rebuilt because of it is not
the forge's business to know.
The forge's module already declares the event. The build machine declares that it consumes nothing.
## What the trigger cannot be
**Not one build per changed module.** The modules form a graph: several are built from one
repository, and some are the base another is compiled on — a runtime image, a compiler base, a
repository whose context a second module builds from. Firing a build for each changed module
independently would start work that cannot succeed yet and produce a failure per dependent, for one
cause.
Observed while catching the mesh up by hand on 2026-09-27: a compiler base had to move before
anything compiled against it could build, and when it failed, the right behaviour was for its
dependents to wait rather than each fail the same way. Fifteen modules shared the cause. A trigger
that reports it fifteen times has buried it.
So whatever reacts to the forge's event resolves what changed into an order, builds the bases first,
and holds a dependent while its base is unbuilt or failed. That is a larger thing than "rebuild what
the commit touched", and knowing it now is cheaper than discovering it from fifteen identical
failures.
## Open questions
- Is "the source moved" still a thing a person can assert by hand once the event path exists, or does
the hand-operated form become the thing that made this failure possible?
- Which commit does the build machine act on — the merge, or each commit it brought — and what does
it do when several arrive for one module at once?
- How does the overview stop being able to lie? Comparing what was built against what was recorded
will always agree. Whether the trunk has moved is a question only the forge can answer, so either
the overview asks it, or it stops claiming to know.
- Does this want to be the same mechanism as the build request on the bus
([ADR 0129](../../02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md)), or does it sit in
front of it?
## What was done (2026-09-28)
The shape above was built as described: the forge's module emits the merge, the control plane
consumes it, and nothing on either side knows the other's internals. The build is asked for the
merge commit, not each commit the merge brought — the trunk moved once, to one place. Several merges
for one module arriving in a row are followed in turn, each moving the recorded source to its own
commit, so the last one to arrive is the one the mesh ends up built from.
The hand-operated form stays. `module moved` is how a source is recorded without a forge — a module
built from a repository elsewhere, or a mesh whose forge module is down — and it is the same act the
event performs, so the two cannot disagree about what "moved" means.
**Bases first needed edges, and there were none.** The order this report asked for was written and
walked a graph that no build had ever recorded: a recipe reads its base from a build argument, so the
digest was never in the file the builder read edges from. A build now reports what it was handed, the
control plane records it by artifact path, and the order is read from each module's newest build.
**What still can lie.** The overview compares what was built against where it was last told the
source is; the forge's event is now what moves that mark, so it is right for as long as the forge
module was listening. A merge made while that module was down is a merge the mesh does not know of
until the module next polls — it announces what merged since it last looked, so the gap closes when
it comes back, and not before. The overview does not say so.
@@ -0,0 +1,60 @@
---
status: resolved
opened: 2026-09-28
located-in: [mesh-controller cmd/mesh-controller]
fixed-by: mesh-controller — `module add` takes `--path` and `--self`, so a module handed over by hand records the whole location it came from; a record naming a repository and no directory says so in the reply; the rule is one function with a test beside it. The nine records already wrong were corrected by rebuilding each with its real directory, which is the same act through the same door.
amended-design:
---
# 132 — A module can be recorded without the directory it lives in
## What was observed
Nine modules on one mesh could not be rebuilt. Each attempt failed the same way:
> has no module.json at its root, so there is nothing saying what it is
All nine were recorded as coming from a repository that holds many modules, each in its own
directory — and each record named the repository and no directory. So every build cloned the
repository and looked for a manifest where there has never been one.
The failure only surfaced when something asked for all of them at once. Before that, the overview
said every module was current with its source, because what it compares is what was built against
what the mesh was last told the source has, and neither half knows whether the source can be found
at all.
## Why it matters beyond this instance
**A module is a repository and a directory inside it** ([ADR 0069](../../02-DECISIONS/0069-a-module-is-a-repository-and-a-path.md)),
and one of the two doors into the catalogue could record only the first half. A build records the
directory it was given, so a module that arrived by being built is always whole; a module handed over
by hand had no way to say where it lived, and the flag to say it did not exist. The rule was decided
and enforced on one path out of two.
**Half a location reads exactly like a whole one.** Nothing in the record is empty in a way a person
would notice: the repository is there, the branch is there, the commit is there. The mesh only finds
out at the moment it needs the manifest, which is the moment it is trying to rebuild — and the module
stays on whatever it last built, indefinitely, with nothing saying why.
**It is the same shape as [131](../131-nothing-tells-the-mesh-a-source-moved/00-report.md).** A
comparison between two facts the mesh holds about itself will agree with itself. Whether the source
can be found is a question only an attempt to read it answers, and the answer had nowhere to go.
## What was done
`module add` takes the directory and which forge holds the repository, so a hand-registered module
records the same whole location a built one does. What a record must say to be worth anything is one
function with a test beside it, rather than a paragraph in a help string: provenance together or not
at all, a directory needs a repository to be inside, a path on the mesh's own forge is not an address.
And a record that names a repository but no directory says so when it is made — not refused, because a
module really at a repository's root is ordinary, but said, because the person adding it is the one
who knows which it is.
The nine wrong records were corrected by building each with its real directory, which re-records it.
No row was written by hand.
## What is still true
A directory that does not exist in the repository cannot be refused when the module is added: the
control plane does not clone, and inventing a check there would mean it did. The first build says so
plainly, which is one build rather than nine, and the record it leaves behind is right from then on.
@@ -0,0 +1,78 @@
---
status: resolved
opened: 2026-09-28
located-in: [mesh-controller module.json]
fixed-by: mesh-controller — the control plane's module declares a run-once `migrate` step before its server, which is the shape ADR 0052 prescribes for exactly this. A step's record of having run is the digest of its declaration and the image is part of that digest, so a new build of the control plane re-runs it; and because a run-once step gates what the declaration places after it, a migration that fails stops the new server from starting at all rather than letting it run against a schema it does not have.
amended-design: 03-DESIGN/01-to-be/32-what-a-module-declares.md
---
# 133 — The control plane's schema is migrated at birth and never again
## What was observed
On 2026-09-28 at 08:17 the control plane was replaced, by the mesh's own upgrade path, with a build
whose code writes a column that a migration **in that same build** creates. Nothing ran the migration.
For the next three quarters of an hour the mesh built things and recorded none of them. Every build
answered:
> ERROR: column "built_contexts" of relation "build" does not exist (SQLSTATE 42703)
and that sentence went only to whoever happened to be waiting on a build's reply. The overview kept
saying the mesh was fine. The builds themselves worked — images were built and published — so the
registry filled up with artifacts the mesh has no record of, and the graph stopped learning without
anything saying so.
The schema was created once, at genesis, by an action in the foundation bundle that runs the same
binary's `migrate`. Nothing runs it again. The mesh has updated its own control plane many times since
that bundle, and every one of those updates carried whatever migrations the new build brought and
applied none of them. This is the first time a build needed one.
## Why it matters beyond this instance
**The schema and the code that needs it ship as one artifact and are applied by two mechanisms, only
one of which is automatic.** A module's version is atomic everywhere else in the mesh — the manifest,
the image and what the machine runs move together. Its schema did not, so "the mesh updates itself on
a push" was true of the code and false of what the code needs.
**The failure is quiet exactly where quiet is worst.** A build that cannot be recorded is a build that
happened and left no trace, which is the fault [issue 050](../050-the-catalogue-knows-nothing-built-before-it/00-report.md)
and [issue 131](../131-nothing-tells-the-mesh-a-source-moved/00-report.md) are both about. The mesh
has three mechanisms for noticing a module is behind its source and none for noticing that what it
recorded was refused.
**The shape was already decided, and the control plane was the one module that did not use it.**
[ADR 0052](../../02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md) says a run-once container is a
step the host runs to completion before whatever the declaration places after it, and names migrating
a schema as the case it exists for. The genesis code's own comment says a manifest may name its image
in more than one resource — "a migrate step beside the server". The control plane's manifest had no
such step; it went straight from a state directory to the server.
## What is still true
**Additive migrations are load-bearing, not a style preference.** The step runs before the *new*
server starts, which means the old binary briefly runs against the new schema. A migration that
removes or renames something would break the running control plane in the window between the two.
**A hand-written step is one the next module forgets**, which is why this fix is not where the matter
ends: [ADR 0135](../../02-DECISIONS/0135-a-module-version-prepares-its-state-before-it-runs.md) makes it
derived and puts it where an author works: a module version declares an entrypoint that prepares its
state, and the mesh composes the gated work from it, so the control plane stops being the only module
that had to remember. That record also settles the level question HAL answered with stages — a consumer
is a module on a machine, so the scope of preparation is the scope of the state — and
[ADR 0134](../../02-DECISIONS/0134-the-mesh-says-what-it-applied.md) answers the second open question
below: what a machine applied, and what it refused, become facts on the bus rather than a line in a log.
**The mesh now has two shapes for one problem.** The catalogue module migrates its own schema in its
own code when it starts; the control plane migrates in a step the host gates on. Both work and the
reasons differ — a module that owns its store entirely can do it at start, while a step is visible in
the declaration and refuses to let a broken upgrade serve. Which one the mesh should standardise on is
a decision, not a fix, and it is not made here.
## Open questions
- Should a module be refusable at registration when it ships migrations and declares no step and no
other way to apply them? The mesh can see both halves.
- Should a record the store refuses reach the overview? Today the only reader of that failure is
whoever asked for the thing that failed, and for an event arriving on the bus there is no such
person.
@@ -0,0 +1,66 @@
---
status: open
opened: 2026-09-28
located-in: [mesh-catalog, mesh-controller internal/catalogue]
fixed-by:
amended-design:
---
# 134 — A definition may still name the mesh, and the check that would say so does not exist
## What was observed
[ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md) says a module
definition names no node, no mesh and no host path, and states how that is checked:
> A catalogue test finds no domain name in any definition value.
There is no such test. Run by hand on 2026-09-28, across the 72 manifests in the catalogue, the
question it asks has 15 answers. They are not all the same kind of thing, and the difference matters
more than the count:
**Values the mesh acts on** — seven:
| module | where | what it names |
|---|---|---|
| keycloak | `env.KC_HOSTNAME` | this installation's public name for itself |
| minio | `env.MINIO_BROWSER_REDIRECT_URL` | the same, for its console |
| invoicing | a resource's `image` | a named registry rather than the mesh's artifact store |
| builder | `build.artifacts[].context.repository` | the forge, by URL |
| route-proxy | `build.artifacts[].context.repository` | the forge, by URL |
| route-adapter | a resource's `content` | a proxy's dynamic configuration |
| novox.be | `module` | the module is named after the domain it serves |
**Prose** — eight, in `listens[].why`: de-spiegel, mailu, n8n, only-office, photos, photos-eef,
photos-filip, portainer. Each explains what a port is for and mentions the public name it is reached
by. Nothing reads these; a check written as a string search would report them, and reporting them as
violations of the same rule would be wrong.
## Why it matters beyond this instance
**An unenforced rule is indistinguishable from a wrong one, and costs more, because people believe
it.** The record says the mesh is name-agnostic, four design documents rest on that, and a reader
checking whether it holds finds that it does not — in the places that matter most. The two forge URLs
are what a build reaches into for its source; the two hostnames are what a service tells a browser
about itself.
**It is the difference between a mesh and this mesh.** A definition carrying `novox.be` is a
definition that can only be installed here. The whole point of the rule is that the same catalogue
raises a different mesh with a different name, and today seven modules would need editing to do it.
**And the shape of the fix is not the same for each.** A public name is an operator's choice about an
assignment, which ADR 0112 already provides for; a forge URL should be a path on the git seat
([ADR 0111](../../02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md)); an image from a named
registry is a question about the artifact store, not about naming. Counting them together would hide
that.
## Open questions
- Does a domain in a `why` string break the rule? It is documentation the mesh never reads, and a
check that cannot tell the two apart will either pass things it should catch or fail things nobody
should change.
- Where does a service's public name live, concretely — a setting on the assignment, or a fact the
mesh composes from the node's domain? ADR 0112 says a requirement the mesh resolves; the two
hostnames above are the first real cases.
- Should a build context name a repository on the git seat rather than by URL, and if so, what does
that mean for a context in *another* mesh's forge?
@@ -0,0 +1,98 @@
---
status: resolved
opened: 2026-09-28
located-in: [mesh-host internal/apply]
fixed-by: mesh-host — a container's mesh names are part of the spec digest the host compares, sorted so the digest does not move for a reordering. A container whose names moved is now recreated like a container whose image moved, and the test fails against the previous behaviour.
amended-design:
---
# 135 — A container's mesh names are not compared, so a moved address is never noticed
## What was observed
One container on this mesh had been restarting every thirty seconds for five days — 2286 times — and
the mesh reported the machine as doing what it was told.
Its logs said its database connected and then a query timed out. The database was reachable: the same
query from the same network, with the same credential, answered in milliseconds. What differed was the
name. Inside that container, `novox.internal` resolved to `10.42.0.1`; in every other container on the
machine it resolved to `10.10.0.1`. The mesh's overlay range had moved, and this container still held
the old one:
```
umami created 2026-09-23 novox.internal:10.42.0.1
mesh-catalog created today novox.internal:10.10.0.1
```
A container resolves other machines and public names through the entries the mesh gives it when it is
created, and nothing re-reads them afterwards. The host compares a container against what was declared
by a digest of its spec — image, name, environment, ports, volumes, arguments, resolver, address, and
what it reads — and **the mesh's names were not in it**. So this container matched what was declared,
was left alone, and kept an address that had not existed for five days.
Forty-eight other containers had current names. Not because anything corrected them: each had been
recreated for some other reason — a new image, a changed file — and picked up the current roster on the
way. This one's image is an upstream release that had not moved, and nothing else about it changed, so
nothing ever recreated it.
## Why it matters beyond this instance
**It is the exact fault [issue 045](../045-a-container-keeps-the-values-it-started-with/00-report.md)
named, in the one field that was left out.** That issue is why the digest carries what a container
reads: "a container whose configuration had since been rewritten compared equal and was left alone —
running values the machine no longer holds, while every check reported success." The same sentence
describes this, with *names* in place of *files*.
**The failure is invisible in exactly the way that matters.** The container runs, so the machine
reports it applied. It restarts, but a restarting container is a normal sight during an upgrade. The
only account of the fault is inside the container's own log, in the words of the application rather
than of the mesh — and what it says is that a query timed out, which points at the database.
**And it is most likely to bite what changes least.** Every container that is rebuilt often repairs
itself by accident. The victim is the module whose image is stable — which is to say, the module that
was working fine.
## What was done
The mesh's names are part of the digest, sorted so the digest does not move for a reordering nobody
made. A container whose names moved is now recreated exactly as one whose image moved.
The first apply after this recreates every container that carries mesh names — one restart each,
already the price the mesh pays for any image update — because their recorded digests predate the
field.
## What is still true
The mesh gives a container its names at creation and has no way to change them in place. That is the
container runtime's shape, not a choice; the answer is to recreate, which is what this does. A module
that would rather re-read a roster from a file can already ask for one as a fact
([ADR 0120](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md)) and restart on
it.
## The container was umami, and it had its own record (2026-09-30)
The container in the table above is umami, and its symptom had already been filed three days earlier as
[issue 118](../118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md) — the store
answering the dial and timing out the query, a restart loop, and a public name answering `502`. Neither
record noticed the other: 118 was reasoning about path-MTU and conntrack on the network between the
container and the store, which is what a stale name looks like from inside the container.
118 is resolved by this record's fix and says so. Noted here because a symptom filed twice, diagnosed
once, and closed in one place is how a repository comes to disagree with itself.
## What replaced this fix (2026-09-30)
The fix here — putting the mesh's names into the digest the host compares, so a container whose names
moved is recreated like one whose image moved — worked, and cost more than it was worth. It made the
roster part of every container's identity, so one name moving replaced every container in the mesh: a
module assigned on one machine restarted the store, the registry, the edge and mail on another
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)).
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) goes at
the cause this record only described: **a copy taken at creation is stale the moment the roster moves,
and detecting that is not as good as not copying.** A container resolves the mesh's names through its
machine's resolver, at the moment it asks, so the fault this record reports cannot occur rather than
being noticed a restart later.
Recorded here because this is where somebody arrives to find out why the digest carries names, and the
answer is that it did, for two days short of a month, and stopped.

Some files were not shown because too many files have changed in this diff Show More