Compare commits

...
Author SHA1 Message Date
jschoubben 0cf1ad5dad Merge pull request 'Issue 172: the ssh client block matches one spelling of a machine's name' (#221) from issue/172-the-ssh-client-block-matches-one-spelling-of-a-machine into main 2026-09-30 13:25:34 +00:00
jschoubben 69a002fce3 Issue 172: the ssh client block matches one spelling of a machine's name
Reported by the operator: ssh by the bare name logs in, by the mesh name
is refused. The predecessor's generator writes the bare name only; the
mesh's ssh-client roster already matches both and is not yet shipped.
2026-09-30 15:25:30 +02:00
jschoubben 16a1a52cd8 Merge pull request 'Issue 171: a module that names its own resolver knows no mesh name' (#220) from issue/171-a-modules-own-resolver-knows-no-mesh-name into main 2026-09-30 13:20:30 +00:00
jschoubben af170e3a67 Issue 171: a module that names its own resolver knows no mesh name
Found and fixed the afternoon ADR 0148 landed: mailu-admin lost its
database behind Mailu's own resolver. Two catalogue PRs; an insight on
0148 that a container's dns is a decision, not a preference.
2026-09-30 15:20:26 +02:00
jschoubben 6c2d5f5913 Merge pull request 'Group 2 is resolved: containers resolve, nothing is copied, a route's name says where it arrives' (#219) from issue/110-resolved into main 2026-09-30 13:07:14 +00:00
jschoubben 04c9500b5b Group 2 is resolved: containers resolve, nothing is copied, a route's name says where it arrives
Issue 110's cause was not the filter: the runtime had never been told,
and the resolver dropped a query arriving on a bridge. ADR 0148 step 3
landed once it did (109, 151 resolved). ADR 0151 composes a route's
internal name under the serving node and drops the suffixed alias
(139, 157 resolved). Design 08 amended; a fact in 0148 corrected.
2026-09-30 14:56:43 +02:00
jschoubben 846c1f85f2 Merge pull request 'Issue 107 is resolved: a declaration carries its order' (#217) from issue/107-resolved into main 2026-09-30 12:13:57 +00:00
jschoubben 9eef0bd525 Issue 107 is resolved: a declaration carries its order
Hosts first, then the controller — a build and a push each, now that the
mesh delivers the host. The host refuses a lower sequence than it kept
and drains a batch by sequence rather than arrival; the controller
numbers each send under the node's hold, inside the signed bytes.

Measured: two pushes, sequence 2 in the kept declaration, counters in
the store agree, no machine reads as behind. That last one is the
subtlety: the mesh compares the digest of what it would send against
what it did, and a number changes the bytes, so the read-only comparison
composes with the last number sent rather than a fresh one.
2026-09-30 14:13:50 +02:00
jschoubben 6e08cdf3d6 Merge pull request 'Every machine self-updates, verified, and 107's gate has opened' (#216) from issue/142-self-update-on-every-machine into main 2026-09-30 11:51:53 +00:00
jschoubben 02f291a129 Every machine self-updates, verified, and 107's gate has opened
All four machines run a host the mesh built, published and delivered, the
last delivery unattended: each stood aside once for a genuinely newer
version and the delivered launcher started it. A following push that
delivered nothing new was applied and reported by every machine and stood
nobody aside.

The crossover needs one restart of the unit per machine, once, because
the running launcher executes from its own inode. Measured timing: three
seconds on the machine, 17-20 as the operator sees it, the difference
being the control plane composing before it sends.

107 is unblocked: a declaration field is now a build and a push.
2026-09-30 13:51:46 +02:00
jschoubben 9b14430d3f Merge pull request 'Issue 163: a delivered host stood aside on every push and reported nothing' (#214) from issue/163-a-delivered-host-stands-aside-on-every-push into main 2026-09-30 11:47:35 +00:00
jschoubben a4384f13d3 Issue 163: a delivered host stood aside on every push and reported nothing
Asked whether a newer host was delivered using the link-time stamp, which
every delivered host carries as 'development build' now that the version
comes from where the binary sits. Never matched, so it stood aside on
every push for ever; standing aside cancels the report, so the mesh never
heard from it. Read as healthy throughout.

The three-minute push wait made it invisible: a wait long enough to
absorb a whole apply is long enough to hide that the machine never
answered.
2026-09-30 13:47:28 +02:00
jschoubben a841e2c173 Merge pull request 'The host self-updates, and an archive cannot be undeclared' (#212) from issue/161-resolved-and-162-an-archive-cannot-be-removed into main 2026-09-30 11:16:44 +00:00
jschoubben 3c535ead31 The host self-updates, and an archive cannot be undeclared
161 resolved and verified on a machine: the workstation runs a host the
mesh compiled, published, delivered and started, applying declarations
and reporting the version it was delivered as.

The system it was built for comes from the artifact — the one thing a
toolchain takes from a module, which 0142 already allowed because the
target is a property of the artifact. The version comes from where the
binary sits, which 0142 decided and nothing had implemented.

Two mistakes on the way, both caught by reading the output rather than
the line that claimed success. A second -ldflags does not merge with the
first: the binary gained its system and lost -s -w, 12.2MB against 8.5MB.
And the delivered binary was named after its package, so the first
delivery was correct, reported success and was invisible to the launcher.

A delivered host that cannot apply is a machine the mesh cannot repair,
because the declaration that would fix it is the one it cannot apply. The
launcher's fallback is what made that an inconvenience instead of an
expedition.

162 is new and not about the host: an archive has no removal, so a module
using one can never be unassigned, and the attempt takes the whole apply
with it — the machine applies nothing else either. It is how undoing the
first delivery froze the workstation.
2026-09-30 13:16:37 +02:00
jschoubben 4cf941d858 Merge pull request 'Self-update works, and a delivered host is one fact short of usable' (#211) from issue/161-a-delivered-host-has-no-link-time-facts into main 2026-09-30 10:28:08 +00:00
jschoubben 5042ffd8d3 Self-update works, and a delivered host is one fact short of usable
The loop closed on the workstation: the version landed, the launcher was
replaced, the running host stood aside, and after one restart the launcher
started a binary the mesh had compiled, published and delivered.

The launcher goes as a file resource rather than inside the archive, and
that is the safety rather than a preference. A file is written atomically,
so the running launcher keeps the inode it started from; an archive writes
in place with truncate and would cut a script a shell is reading. The
manifest carries a second copy and a test refuses any drift from the one
in packaging.

Then it would have refused the first declaration it was asked to apply.
The Makefile links in two facts the mesh's toolchain does not, on purpose,
and one of them is the system the host was built for — read before
anything is applied, so the failure is safe and total. Nothing reports it:
the unit is active, the bus link is up, and the log says it is hearing
what the node should be.

Worse, the declaration that would fix it is the declaration it cannot
apply, so the mesh cannot repair such a machine. Restored by moving the
delivered versions aside and letting the launcher fall back, which is the
fallback working as designed.

0142 already settles the version — it comes from where the component sits,
not from its linker — and that is unimplemented. The system pin has no
answer, and the candidates are a decision rather than a fix: put it in the
path too, carry it in a file beside the binary, or stop pinning at link
time at all, which is 0005's to change.
2026-09-30 12:28:01 +02:00
jschoubben 4f9dc3906e Merge pull request 'A machine says little about itself, and only when the mesh asks it something' (#209) from issue/160-what-a-machine-says-about-itself into main 2026-09-30 08:36:48 +00:00
jschoubben a2542e51f8 A machine says little about itself, and only when the mesh asks it something
Filed as a to-do. Nothing is broken by it: every machine here is amd64 and
reports so, and one architecture is enough for now.

When a machine joins, the mesh should collect what it reasonably can about
it and refresh that daily. It already asks what a machine can do; what it
is made of is the same question one level down.

More is already collected than it looks — eight capabilities, the links
that face outside, the host version, and on an adopted machine what it
holds, what is reachable and the firewall and tunnel it was found with.
The architecture and the kernel are in there too, and nothing reads
either: measured, all four machines report amd64 and linux, and node show
prints the capabilities beside them without printing them.

Missing: memory, disk, the processor beyond its architecture, the
distribution and its version, virtual or physical, cores, uptime. Several
are what somebody wants when deciding where a module goes, and the
placement code's own comment already imagines them.

Also missing: the refresh. A machine publishes after an apply, and the
five-minute reconcile publishes nothing, so the mesh's picture is as old
as the last push. Same mechanism 087 wanted.

This is the third thing in one day found to be collected and read nowhere,
after held resources and the host version. Whatever gets added should say
in the same breath which surface shows it, or it will be the fourth.

159 gains the note that the architecture is already reported, so matching
an artifact to a machine needs no new fact — only the comparison and a
compiler told what to target.
2026-09-30 10:36:40 +02:00
jschoubben b19b29cd3a Merge pull request 'An artifact's system is checked and then nothing uses it' (#208) from issue/159-an-artifacts-system-is-checked-and-ignored into main 2026-09-30 08:30:52 +00:00
jschoubben 4bf4fe2d06 An artifact's system is checked and then nothing uses it
Asked whether the host is built for more than one architecture. It is
not, and the reason is worse than a missing feature.

A bundle in a compiled language must name a system, must name one of
alpine, android or arch, and is refused with a careful message if it gets
that wrong. The field is then read by nothing: it does not reach the
compiler, no machine is matched against it, and nothing chooses between
two artifacts by it. The compile runs with no target named and produces a
binary for whatever the build machine happens to be.

The host is x86-64 because the build machine is, not because the
declaration said so. Correct for this mesh by coincidence — four machines,
all x86-64 Arch.

A module declaring two systems would get two identical binaries, both
published and both pinned, and the one sent to the machine it was not
built for would fail at exec. android is the sharp end: not an x86-64
platform, and an artifact declared for it today would be an x86-64 binary
wearing the label.

A field that is checked and ignored is worse than one that does not
exist, because the check is what persuades you it works.

Also noted: the processor is a second dimension the manifest has no word
for, so even implementing the present field would not answer the question
that found this. And since the Go toolchain builds statically, one binary
would run on all three systems anyway — so the pin is a policy rather
than a necessity, which is a decision and not a fix.
2026-09-30 10:30:45 +02:00
jschoubben 5e12081774 Merge pull request 'Issue 142: the mesh compiles its own host and publishes it to its registry' (#207) from issue/142-the-mesh-can-build-its-own-host into main 2026-09-30 08:05:12 +00:00
jschoubben 9ad64881a9 Issue 142: the mesh compiles its own host and publishes it to its registry
Both of the things ADR 0141's insight named as remaining are built. A Go
toolchain based on a new mesh-tools-go module, so the compiler is named
and not pinned; and ${version} in any value of a resource that uses an
archive or a bundle.

Measured rather than asserted: the mesh built the host through its own
toolchain, published it to its own registry, and the bundle fetched back
out is a statically linked stripped binary that runs and says it is the
host.

The cost was larger again than 0141's note said. Three more things in the
path assumed one language or one shape — an entrypoint became a .ts file
whatever the language, the output directory was the compiler's to create,
and a bundle was refused if it named what it is built from — and a fourth
was in the base image, which is Alpine where the first Dockerfile ran
apt-get. That last one is issue 136 in an image, and the build refused
rather than a module failing later.

The version in a path is the digest, not the commit: two builds of one
commit are the same bytes, so a content-addressed version keeps the path
an unchanged build already had.

Still nothing delivers a version to a machine. The host module declares no
resources, so the bundle sits in the registry and no machine is asked to
take it. 0141 carries the insight and 142 the account.
2026-09-30 10:05:04 +02:00
jschoubben 266ade6e28 Merge pull request 'Issue 087: what I shipped first said the opposite of the truth' (#206) from issue/087-a-commit-has-no-order into main 2026-09-30 07:29:53 +00:00
jschoubben a794ef9a3f Issue 087: what I shipped first said the opposite of the truth
It reported "N machines run an older host than another" by comparing
versions as strings. A host reports its version as a commit, and commits
have no order. On the live mesh it named the three machines running the
NEWER host as the ones behind — ced54d4 sorts above 04a27ca and that is
all it means.

The code even carried a caveat saying versions compare as strings and that
this "is enough for the timestamps and commits this mesh uses". That was
the error, written down and not noticed: enough for timestamps, meaningless
for commits, and the mesh reports commits.

It now reports the split and claims no ordering, which is more useful as
well as more honest — the reader sees who is on which side, and that is
what decides whether a field can be sent. Ordering is left with the host,
which would have to report something ordered for anybody to have it.

This is issue 145 arriving by my own door an hour after I closed it: a
report that confidently says the opposite of the truth is worse than one
that says less.
2026-09-30 09:29:40 +02:00
jschoubben e10084fed6 Merge pull request 'Issue 087: a machine states its host only when the mesh sends it something' (#205) from issue/087-a-machine-says-its-host-only-when-asked into main 2026-09-30 07:09:17 +00:00
jschoubben 941b920bd7 Issue 087: a machine states its host only when the mesh sends it something
Measured live. All four machines run the identical host binary — same
digest, installed within eighteen seconds — and at first only one reported
a version, which read as a difference where there was none.

A machine publishes a report after an apply. The five-minute reconcile
publishes nothing, because it is the machine keeping itself as declared
rather than answering anything. So a current, idle machine never says, and
the mesh cannot tell that from a machine running something ancient.
Confirmed by pushing: not reported, then 04a27ca.

Enough for the purpose, not enough for the claim. For deciding whether a
new declaration field is safe it is sufficient — pushing is what the mesh
is about to do, and the answer arrives with the act. For knowing what the
mesh runs it is not, and "not reported" is worded as "nobody has asked
recently" for that reason.

Making a heartbeat carry it would close the gap and would change what a
heartbeat is — a bare word that the node is there, deliberately carrying
nothing else. Left alone rather than widened in passing.
2026-09-30 09:09:10 +02:00
jschoubben ecbd2ce2e6 Merge pull request 'Group 1: 145's report states its scope, and 107 waits for delivery' (#204) from issue/145-and-107-what-group-one-leaves into main 2026-09-30 07:00:32 +00:00
jschoubben b7f7b97d8a Group 1: 145's report states its scope, and 107 waits for delivery
145, partly resolved. The sentence that was true for eleven hours of a
mesh in which no module could reach another now says what it is not a
claim about: that is the mesh and the machines agreeing, and nothing here
dials a provision. It checks nothing and does not pretend to — ADR 0146
decides the check and is deliberately not built. What changed is that the
report no longer implies otherwise. Stays open for that reason.

Carried forward: 0146's check needs an internal name fetched over TLS with
the certificate verified, and until today no machine trusted the mesh's
authority. Three of four do now, so whoever builds it does not have to
solve that first.

107, diagnosed and deliberately not built. The premise is confirmed in the
host's own words — unknown fields are refused because "a field the host
does not know is a thing the control plane believes it asked for" — so the
fix is a flag day, not an addition. 087 now makes the cost measurable, and
the measurement is why it waits: one machine of four runs an older host,
it is parked, and nothing delivers a host at all (142). Shipping the field
means hand-placing binaries and unparking a machine, and one missed in
that sequence is unreachable, not degraded. The fault it prevents has
never been observed.

142 gains the note that it is 107's gate, and that it is what makes a
declaration field cost a rollout instead of an expedition.

A judgement about order, not a refusal, and cheap to overrule.
2026-09-30 09:00:18 +02:00
jschoubben c2cf0d72d6 Merge pull request 'Issue 087 is resolved: the mesh knows which host runs a machine' (#203) from issue/087-the-mesh-knows-which-host-runs-a-machine into main 2026-09-30 06:55:05 +00:00
jschoubben 0a5006b366 Issue 087 is resolved: the mesh knows which host runs a machine
The machine has reported its host version since ADR 0141, whose own
comment says why it must: without it nothing can say a machine is behind.
The controller's copy of the report did not have the field, so it
unmarshalled into nothing and was thrown away on arrival. Two structs
describe one message and only the sending side had it.

node show names it per machine, "not reported" where the mesh has not been
told. status names every machine running an older host than another does,
and which is newest.

Disagreement rather than staleness, deliberately: nothing delivers a host
version yet, so the mesh holds no canonical current one and "behind" has
no fixed point. What it can say is that the oldest host in the mesh is
what the mesh may send.

Two refusals to guess: a machine that reported nothing is not called
behind, and versions compare as strings — right for the timestamps this
mesh uses, wrong for a scheme where 10 sorts before 9, said at the place
that would have to learn.

107 gains the note that this is what makes its new field safe to consider,
and that one machine of four is behind today, so it is not free yet.
2026-09-30 08:54:52 +02:00
jschoubben 3ea9601904 Merge pull request 'Issue 125 is resolved: a hold is a line in the report, and it breaks "all well"' (#202) from issue/125-a-hold-is-a-line-in-the-report into main 2026-09-30 06:47:57 +00:00
jschoubben f7a37ee4f3 Issue 125 is resolved: a hold is a line in the report, and it breaks "all well"
Two of the four surfaces the report named already carried it — the host
has reported Held since ADR 0100, and node show reads the machine's own
list with an `as of` beside it. Recorded as checked rather than assumed.

Two did not. The apply line counted what it applied and said nothing
about the difference; status read the mesh's take-time listing, so a
module assigned after it showed nothing at all.

Both now say it, and the part that carries the weight: a hold suppresses
"all doing what they were told, all heard from, running what the mesh
would send them". That sentence was true for the whole outage, and acting
on it is what stopped the predecessor's proxy. Being adopted still does
not suppress it — a mode somebody chose is not a half-finished action.

Status does not call a hold a fault, deliberately. It is correct
behaviour, and a reader trained to see red for something the mesh did
right stops reading.
2026-09-30 08:47:36 +02:00
jschoubben 601d004fcb Merge pull request 'Issue 129: three machines of four trust the mesh, not one' (#201) from issue/129-three-machines-not-one into main 2026-09-30 00:30:30 +00:00
jschoubben d009c3efef Issue 129: three machines of four trust the mesh, not one
Extended after the first was proven. Every converged machine now holds
the anchor and verifies an internal name with a plain client; before,
the two unassigned ones answered 'unable to get local issuer
certificate' and held no entry for the mesh.

ace is excluded on purpose: it is adopted, so a module assigned there is
held rather than run, which is right and is not trust.

Both the resolution and 0147's insight said one machine of four, which
was true for about twenty minutes.
2026-09-30 02:30:23 +02:00
jschoubben 5036b927b9 Merge pull request 'Issue 129 is resolved: a workstation trusts the mesh, and stops when told to' (#200) from issue/129-a-machine-trusts-the-mesh into main 2026-09-30 00:18:51 +00:00
jschoubben 1f72e82b84 Issue 129 is resolved: a workstation trusts the mesh, and stops when told to
Registered ca-trust from the catalogue — it was merged and had never been
registered, which is why "assign it to one machine" had no module to name
— assigned it to the workstation, and verified.

Verified in the form ADR 0147 prescribes, against the authority's own API
so the handshake needs nothing else in the mesh to be right: 200, issuer
Mesh Internal CA, Verify return code 0. Four routed internal names verify
too, and `git ls-remote https://…` works, which is the consequence the
report named.

Removal exercised for the first time. Unassign and push removes the
anchor, empties the trust store of the mesh's authority, and returns the
plain client to the original error; assigning again restores it. That is
the half 0147 claimed and nothing had shown.

One thing it found that is not in the module: removal works only because
the host removes the service before the script. Stopping the unit is what
deletes the certificate and refreshes the bundles, and it needs the script
to still exist. The symmetry rests on an ordering nothing states.

0147's "written, and not yet run" now carries a progressive insight
saying it has run, where, and that it ran on one machine of four.
2026-09-30 02:18:21 +02:00
jschoubben 76fbe323ea Merge pull request 'Issue 118 is resolved: it was issue 135, and umami is healthy' (#199) from issue/118-is-135-and-is-resolved into main 2026-09-29 23:54:39 +00:00
jschoubben 295bf2f3ab Issue 118 is resolved: it was issue 135, and umami is healthy
Verified on the machine: umami has zero restarts, applies all 26
migrations, reports the store up to date, and umami.novox.be answers 200
where it had answered 502 since 2026-09-25.

It was the same fault as 135, filed three days earlier and diagnosed
without either record noticing the other. 135's container IS umami — it
is named in 135's own evidence table, holding novox.internal:10.42.0.1
after the overlay range moved. The dial appeared to succeed and the first
real query timed out because the name pointed at an address that no
longer existed; the store, the path and the credential were all fine.

118's own reasoning is kept as a warning, because it is careful and
wrong: it argued path-MTU and conntrack, and it named the move that would
have found it — compare umami against a container created that week —
and did not make it. A stale name presents as a network fault, which is
why ADR 0148 stops copying names into containers rather than detecting
when the copies go bad.

Also: 118's located-in named mesh-catalog modules/umami, which was never
at fault; it names mesh-host's comparison now. And 135 gains the pointer
to 118, which is the back-reference I have now missed three times.
2026-09-30 01:54:19 +02:00
jschoubben e1b74810a8 Merge pull request 'Issue 129 is live and reproduced, and needs three steps rather than one' (#198) from issue/129-and-what-reproducing-it-found into main 2026-09-29 23:30:48 +00:00
jschoubben e41eed0852 Issue 129 is live and reproduced, and needs three steps rather than one
The certificate is genuine, from Mesh Internal CA, and nothing on the
workstation trusts it — verbatim the error the report gives. The public
name on the same proxy verifies cleanly, which puts the fault exactly
where the report puts it.

What is in the way is not an assignment. `ca-trust` is merged in the
catalogue and has never been registered with the mesh — 39 of 76
manifests are — so there is no module to assign. It dry-runs clean and
needs no artifact built.

Two findings from reproducing it, both their own issues:

157 — every routed name is published with an `.internal` alias that
nothing serves. The hosts file says keycloak.novox.be.internal; the proxy
serves keycloak.novox.internal and refuses the other by name. The first
three names I tried came from the hosts file and failed with a TLS alert
rather than a verification error, which pointed at a regression that had
not happened.

158 — the proxy re-logs all 52 routes every two seconds, 31 times a
minute. The one line that explained 157 sat between two of them.

Also recorded, because it was nearly filed as a defect and is not one:
step-ca publishes roots as /roots.pem, which is PEM, so ca-trust's fetch
and its refuse-a-non-certificate guard are both right. Its other endpoint
/roots returns JSON that contains the text the guard greps for, so the
guard is sound only because of which path is published.
2026-09-30 01:30:24 +02:00
jschoubben ba30286896 Merge pull request 'The pointers back from what yesterday's records changed, which I missed twice' (#197) from decision/the-pointers-back-from-what-these-narrow into main 2026-09-29 22:47:45 +00:00
jschoubben 53b94c51bb The pointers back from what yesterday's records changed, which I missed twice
Three records were left describing a mechanism a new record had moved,
and a reader arrives at them by following a citation: 0066 still said a
routed name is written into every container after 0148 replaced that with
resolution; 0016 still read as though the lab were the test bed after
0149; and issues 109 and 135 said nothing about 0148 ending the copying
that 135's own fix made comparable. Each was a citation leading to the
wrong answer in a record that was not wrong about anything it decided.

This is the second time in one session. The playbook rule I added last
round did not stop it, so the convention is now written where the record
conventions live, with the shape to use and three worked examples — and
with the honest note that it is NOT machine-checked and cannot be from
`extends:` alone: 102 records extend another, 87 have no back-reference,
and that is correct, because extending usually means building on a
context. Making it mechanical means a record declaring the relationship
in frontmatter, which is a schema change and is not mine to decide.

Also: designs 18 and 20 claimed `updated:` dates from before I edited
them, and 117's `fixed-by` gained the commit beside the record.
2026-09-30 00:47:19 +02:00
jschoubben 83791f0921 Merge pull request 'The four open design questions, answered: ADRs 0148, 0149, 0150, and 0114 accepted' (#196) from decision/0148-the-meshs-names-are-resolved-not-copied into main 2026-09-29 22:39:38 +00:00
jschoubben bf4d4e2e7b The three parked questions are answered: 0114 accepted, 0068 superseded, 117 settled
**0114 accepted.** The separation it draws — a consumer's resource and a
credential that reaches it are different things — is a data-loss rule,
and a record naming one should not sit unresolved while the code that
could hit it is written. Not built, and accepting it schedules nothing;
accepted-and-not-built is where 0141 and 0142 already are.

**0068 superseded by 0149, the live mesh is the test bed.** Never built,
and contradicted by practice that was written down nowhere but a handoff
note. The faults that cost the most are faults of a mesh that already
exists — bound consumers, containers made against an older roster, an
adopted machine — and a bed is by construction a mesh that does not.
Issue 156 settles it: that change was verified honestly, on the only path
where it cannot fail. The lab is not retired; raising a mesh from bare is
now its whole job.

**117 answered by 0150.** A module's own code runs as supervised
processes under the module's one account. 0047's argument was for a
runtime per module, not for a container — a unit satisfies it and runs as
an account. The invariant is the account, not the process count: its
worry was a second identity to scope and seal, and processes sharing one
account create none. Designs 18 and 20 now cite it and 0047 carries a
dated pointer, which closes the third disagreement — that neither design
knew the record existed.
2026-09-30 00:39:09 +02:00
jschoubben ec42ee0846 ADR 0148: the mesh's names are resolved, not copied into every container
Answers issue 151. Copying the roster into each container made the roster
part of each container's identity, so one name moving replaced every
container in the mesh — and it never stopped the staleness it was for,
since a copy taken at creation is stale the moment the roster moves (109,
135).

A container resolves through its machine's resolver instead, and nothing
is copied. Staleness stops being possible rather than detected, and a
name's blast radius becomes nothing.

Scoping each container to the names it binds was the close call and is
rejected: it contradicts anything-calls-anything, and leaves the roster
in the digest so the churn returns for a widely-bound name.

Gated on issue 110 — a container on the runtime's default network has no
DNS at all today. Removing the copy first reintroduces 109 and 135
silently on a live mesh. 151 stays open until the code lands; design 08's
file-not-resolver passage is narrowed to the machine's own roster.
2026-09-30 00:34:30 +02:00
jschoubben 5a3dee9e9e Merge pull request 'The records pointed at branches that no longer exist, and two fixes had no sequel' (#195) from issue/pointers-that-resolve-to-nothing into main 2026-09-29 22:29:02 +00:00
jschoubben 47909c6b71 The records pointed at branches that no longer exist, and two fixes had no sequel
Three issues named the branch that fixed them, and a branch is deleted
when it merges — so every `fixed-by:` was a pointer that resolved to
nothing by the time anyone followed it. They name commits and pull
requests now, and playbook 03 says to.

Two records were missing the thing a reader arrives for. 146 did not say
that one of its fixes crash-looped the control plane on a running mesh,
which is the whole reason the delivery subject carries the stream and the
raise path was the only one exercised. 151 did not say that 152 removed
the false reasons its roster moved, or that it stays open for the real
ones.

ADR 0080 enumerates what cycle.py enforces and named four things; it
enforces five. A progressive insight names the fifth — the decision
stands, the list had gone stale. The checks README and playbook 03 gained
the same rule, and 155 points at all three.
2026-09-30 00:28:35 +02:00
jschoubben 7e5edab8da Merge pull request 'Issue 156: the wider grant has landed, and the notes the fix printed' (#194) from issue/156-the-grant-has-since-landed into main 2026-09-29 22:19:35 +00:00
jschoubben 3649f82204 Issue 156: the wider grant has landed, and the notes it printed
Measured on the broker's own config after the fixed controller composed
it: every node is now allowed _DELIVER.<node>.> as well as the bare
subject. That grant was missing when this was diagnosed, which is why
re-making the consumers would have silenced every machine.

Also records the five consumers it kept and reported, including the build
machine's worker, which is bound the same way and was not anticipated.
2026-09-30 00:18:57 +02:00
jschoubben b400ce2c54 Merge pull request 'Issue 156: moving a consumer's delivery subject stops a running mesh' (#193) from issue/156-a-consumer-that-works-is-not-replaced into main 2026-09-29 21:57:00 +00:00
jschoubben 995c8cb266 Issue 156: moving a consumer's delivery subject stops a running mesh
146's change is correct on a mesh being raised and fatal on one that is
running: the server will not move a push consumer's delivery subject
while a subscriber is bound, and a node is bound to its declaration
consumer the whole time it is up. Reproduced before fixing; an earlier
version of the test unsubscribed first and passed against the broken
code.

The consumer that works is kept and the assertion says so. Not re-made:
the nodes are granted the bare subject only, so re-making would have
silenced every machine.
2026-09-29 23:56:31 +02:00
jschoubben f89aef992d Merge pull request 'Two records shared a number, twice, and every check passed' (#192) from issue/two-records-share-a-number-and-nothing-says-so into main 2026-09-29 21:39:09 +00:00
jschoubben b967ef7be3 Two records shared a number, twice, and every check passed
Two machines filing issues in the same hour both read main correctly and
both took "the next free number". main lags every open pull request —
seven that evening — so they collided twice. The second collision
reached main with records, cycle and index all reporting success.

cycle.py now refuses a tree where two issue folders share a leading
number, and names both. Proven by adding a duplicate and watching it
fail. The colliding records become 153 and 154, renumbered in the branch
that lands last, because renumbering a branch whose author is still
pushing only moves the race.

The check catches a collision; it does not prevent one. Taking a number
still means reading the open pull requests as well as main — issue 155
says so.
2026-09-29 23:38:48 +02:00
jschoubben e417906241 Merge pull request 'Issue 150: a machine's own network is not a reach' (#190) from issue/150-a-machines-own-network-is-not-a-reach into main 2026-09-29 21:37:41 +00:00
jschoubben b1bf895688 Merge pull request 'Issue 149: an adopted machine's data cannot be placed where it is' (#189) from issue/149-adopted-data-cannot-be-placed-where-it-is into main 2026-09-29 21:37:34 +00:00
jschoubben 0d64677c70 Merge pull request 'Issue 152: a node the mesh could not read withdrew its names from every machine' (#191) from fix/152-a-lookup-failure-is-not-an-absence into main 2026-09-29 21:37:16 +00:00
jschoubben 04205c1dc8 Merge pull request 'Issues 147 and 148: a route before its module is taken; a new name recreates every container' (#188) from issue/147-148-found-migrating-ace into main 2026-09-29 21:37:09 +00:00
jschoubben 9dc49cd831 Merge pull request 'Grooming: five issues were fixed and never closed, and one is not' (#187) from grooming/stale-issues into main 2026-09-29 21:37:04 +00:00
jschoubben 66b413a076 Merge pull request 'Issue 140 is resolved, and was resolved before it was read again' (#186) from issue/140-resolved into main 2026-09-29 21:36:57 +00:00
jschoubben fcba05fed9 Merge pull request 'Issue 146: a first node now enrols, and is enrolled twice' (#185) from issue/146-diagnosis into main 2026-09-29 21:36:42 +00:00
jschoubben 18f37c25b2 Issue 152: the loop is metastable, and it cleared at 23:11
It ran 22:46-23:11, five full replacements, and stopped when a pass
happened to read the store during a window it was up. The record said
nothing about it stops; that was wrong. Exiting by luck is the finding,
not a mitigation.
2026-09-29 23:27:51 +02:00
jschoubben 3c2b4fc6b6 Issue 152 is fixed: a lookup failure is no longer an absence
The three gatherers pass over a node whose set does not compose, and
raise anything else. 151 stays open: this removes the false reasons a
roster changes, not the fact that a real change still replaces every
container.
2026-09-29 23:26:20 +02:00
jschoubben 72eaf52867 Issue 152: a node whose plan will not compose withdraws its names from every machine
ace numbered its two records 147 and 148, which this repository already
uses; they become 150 and 151, as 127 became 149.

The new record is why the control node could not stop applying: the
roster alternates between two values because routeNamesInTheMesh
swallows a per-node plan failure, and the roster is part of every
container's identity. The loop closes through the control plane's own
store, which each pass replaces.
2026-09-29 23:13:54 +02:00
jschoubben 741625e725 Merge remote-tracking branch 'origin/issue/147-148-found-migrating-ace' into issue/the-roster-flicker-recreates-every-container 2026-09-29 23:12:57 +02:00
jschoubben c3730b9a23 Issue 150: a machine's own network is not a reach
ADR 0138's reach is internal, public or both. ace's LAN devices (an IoT
switch on mosquitto, unifi access points, plex clients) are none of them;
the only reach that admits them is public, which states the wrong thing and
opens the port to the internet on any machine with a public address.
2026-09-29 23:06:08 +02:00
jschoubben 3bd6f34de3 Issues 147 and 148: a route before its module is taken; a new name recreates every container
Both found migrating the first module on ace (searxng), adopted with route-adapter:
147 — assign is not a no-op for a routed module: its route reaches the predecessor's
proxy while its container is held, and the name serves a dead backend until take.
148 — every push that moved the mesh's names replaced every container on novox, twice,
the control plane's own store included.
2026-09-29 23:01:51 +02:00
jschoubben 2b5119ecd2 Issue 114 is answered: the controller is a process, by ADR 0142
It asked a question rather than reporting a defect, and the question was taken
five days later — the mesh's own components are binaries on the machine, and
third-party software stays a container because an image is the right way to
carry somebody else's build. The delivery of them is issue 142.
2026-09-29 22:38:11 +02:00
jschoubben 96bdffa9bc Two records were numbered 127; the second becomes 149, and is resolved
A number identifies a record, and two were given 127 on 2026-09-27. The one
three documents and three source files cite by number keeps it; the other
becomes 149, says so in its own heading, and its two inbound references are
repointed.

It is also resolved: an empty declaration is sent carrying owns_nothing rather
than skipped, and the host refuses an empty body that does not carry it, so
emptiness cannot be read as a truncated declaration.
2026-09-29 22:33:40 +02:00
jschoubben 4eb16f1028 Three proposed records were already built; two are still yours to call
0037 (where a module lives), 0113 (the vault makes every secret) and 0115 (one
assignment of a module per node) described arrangements the mesh has — a
catalogue of other people's software plus a table filled by module add, a vault
answering a secret requirement six modules make, and a rule the assignment
table's primary key already enforces. Each is accepted against what was built,
and says so in its own words.

0037's other half is not built: a manifest outside this catalogue has no check,
which is issue 148.

0068 (the lab takes requests) and 0114 (a shared credential rotates over two
credentials) stay proposed. Neither is built, and both are decisions rather
than records of something that happened.
2026-09-29 22:31:55 +02:00
jschoubben d199de40db Merge the 146/147 records, which 006's note links to 2026-09-29 22:31:46 +02:00
jschoubben 9a1dc4665c Grooming: issue 006's knowledge base is the predecessor's, and is gone
The record catches README claiming this repository is indexed into a knowledge
base. Nothing indexed it, and since the cut-over there is nothing to index it
into — the surface that answered is on the transport the mesh removed (issue
147). Noted where the record is, so the next reader does not go looking for a
search that cannot exist.
2026-09-29 22:26:13 +02:00
jschoubben 14be8576f8 Grooming: five issues were fixed and never closed, and one is not
088, 089, 120, 128 and 130 each name a commit that is on main and cites them —
the forge's address following a moved port, a route naming its endpoint, a
provisioner asking the backend what is there, the hosts file written into a
marked block, and undeclaring giving a unit back the state it was found in.
Each says it was closed by reading commits rather than by a run, so nobody
reads a green that was never measured.

129 stays located on purpose: ca-trust is merged and no machine holds it, so
the symptom it opened on is still true everywhere.
2026-09-29 22:25:25 +02:00
jschoubben ec8676c225 Issue 140 is resolved, and was resolved before it was read again
Written 2026-09-28, answered the same week by mesh-controller bdf965d and
c68d3a7, and never closed — so the mesh's own account said reach was declared
nowhere while three readers were reading it: the filter, the proxy's names, and
each authority's host policy. The resolution names them and what checks each.

What is left is retiring the older per-port keys the block replaces, which is
not a gap in what reach can say.
2026-09-29 22:19:24 +02:00
jschoubben 3c0f7082e6 Issue 147: the tool surface is not the mesh's, it is the predecessor's
Diagnosed from the configuration: the tool server is HAL's brain, a local
process on the workstation with the predecessor's broker URL in the assistant's
own config. No manifest, no assignment, no seat, no account. Nothing regressed
— the mesh removed a transport this program still dials, and the program was
never part of the mesh. The mesh has a tool model and nothing publishes an
operator-facing surface onto it.
2026-09-29 21:50:50 +02:00
jschoubben 0dd00e88b6 Issue 147: the operator's tools still dial the bus that was removed
Every tool call fails with 'AMQP not connected', on every node including the
local one, because the tool surface still opens an AMQP connection and that
transport was deleted at the cut-over. The mesh reports healthy throughout —
what broke is the thing standing outside asking it questions, so nothing the
mesh checks is about it.
2026-09-29 21:39:55 +02:00
jschoubben eef54917ec Issue 146: the double enrolment was two consumers sharing a delivery subject
Not about enrolment. A push consumer delivers onto an ordinary subject and
everything subscribed to it gets a copy; the controller's two consumers were
both named after it, so both were given the same subject and the one process
acted on every message twice. Enrolment is where it drew blood because a second
enrolment mints a second credential.
2026-09-29 21:32:52 +02:00
jschoubben e9b1010bc0 Issue 146: what made it slow, and what was changed so it is not 2026-09-29 17:45:25 +02:00
jschoubben f6ed3545b7 Issue 146: a first node now enrols, and is enrolled twice
Two more faults behind the three already fixed. The account a token is the
password of was never recorded, and the comment above the issuing code said
it was; issuing now records it. Placing the composed list at genesis is the
other half — the control plane says what it composed and whoever raises the
machine writes it beside the bus, because no declaration can reach a machine
that has not enrolled.

With that a first node enrols. It is then enrolled twice from one attempt,
each minting a credential, and it keeps the answer to the first while the mesh
keeps the second. The trail for that one stops at a duplicate that survived
message-id deduplication.
2026-09-29 17:36:40 +02:00
76 changed files with 3388 additions and 86 deletions
+3 -1
View File
@@ -54,4 +54,6 @@ whose failure has never been observed is a guess about its own correctness.
The development cycle, checked ([ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md)): The development cycle, checked ([ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md)):
a to-be design names a decision, an in-progress/implemented design names its owning code, a a to-be design names a decision, an in-progress/implemented design names its owning code, a
located/fixed issue names its owner, a fixed/resolved issue says what fixed it, a graduated located/fixed issue names its owner, a fixed/resolved issue says what fixed it, a graduated
research overview says what it became. `python3 00-META/checks/cycle.py` research overview says what it became, and no two issue records share a number (issue 155 — the
number is how a record is cited, and `main` lags every open pull request, so two people reading it
allocate the same one). `python3 00-META/checks/cycle.py`
+23 -1
View File
@@ -14,7 +14,8 @@ What is enforced:
its owning code (`code:`) -- no development without a design that says where. its owning code (`code:`) -- no development without a design that says where.
issues a known `status:`; once `located`, `located-in:` names the owner; issues a known `status:`; once `located`, `located-in:` names the owner;
once `resolved`, `fixed-by:` says what fixed it (prose counts -- once `resolved`, `fixed-by:` says what fixed it (prose counts --
"nothing, the capability existed" is an answer). "nothing, the capability existed" is an answer). And no two records share a
number -- the number is how a record is cited.
research a known `status:`; a `graduated` overview says what it `became:`, and every research a known `status:`; a `graduated` overview says what it `became:`, and every
target it names exists. target it names exists.
decisions every accepted record is REACHABLE from the cycle: cited by a design doc's decisions every accepted record is REACHABLE from the cycle: cited by a design doc's
@@ -109,6 +110,27 @@ def main():
"without a design that says where" % status) "without a design that says where" % status)
# ---- issues ------------------------------------------------------------------------ # ---- issues ------------------------------------------------------------------------
# Two records may not share a number. Numbers are taken as "next free after main", and work
# sits on unmerged branches for days -- so two people reading the same main allocate the same
# number, and nothing said so. It happened twice in one evening between two machines, and the
# second collision landed on main with all three checks passing (issue 155). An issue number is
# how every other record cites this one; two records answering to it means a pointer that
# resolves to whichever the reader happened to open.
seen = {}
for folder in sorted(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", ""))):
name = os.path.basename(os.path.normpath(folder))
number = name.split("-", 1)[0]
if not number.isdigit():
continue
if number in seen:
bad(os.path.join("04-ISSUES", name),
"is numbered %s, and so is %s -- an issue number is how it is cited, and two "
"records answering to one means a citation that resolves to whichever the reader "
"opened. Take the next free number across main AND every open pull request"
% (number, seen[number]))
else:
seen[number] = name
for path in sorted(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", "00-report.md"))): for path in sorted(glob.glob(os.path.join(ROOT, "04-ISSUES", "*", "00-report.md"))):
front = frontmatter(path) front = frontmatter(path)
if front is None: if front is None:
+14 -1
View File
@@ -21,7 +21,12 @@ incident someone must **clear**.
## Steps ## Steps
1. Take the next free number. Create `04-ISSUES/NNN-short-name/00-report.md`: 1. Take the next free number — **across `main` and every open pull request**, not `main` alone.
Work sits on unmerged branches for days, so two people both reading `main` allocate the same
number; it happened twice in one hour between two machines, and the second collision reached
`main` with every check passing (issue 155). `cycle.py` now refuses two records sharing a number,
which catches a collision but does not prevent one. Create
`04-ISSUES/NNN-short-name/00-report.md`:
```yaml ```yaml
--- ---
@@ -42,6 +47,14 @@ incident someone must **clear**.
## Rules ## Rules
- Closed issues are never deleted — they are the mesh's symptom-to-component memory. - Closed issues are never deleted — they are the mesh's symptom-to-component memory.
- `fixed-by:` names something that will still exist: a commit or a pull request, never a branch. A
branch is deleted when it merges, so a branch name there is a pointer that resolves to nothing by
the time anybody follows it.
- A fix that turns out to have broken something else is written back into the record that asked for
it, pointing at the new issue. Somebody arriving at a record to learn why the code is the way it
is must not have to already know there was a sequel.
- Renumbering a collision happens once, in the branch that lands last. Renumbering a branch whose
author is still pushing only moves the race.
- An issue whose answer is a general lesson should also be written to the knowledge base, so - An issue whose answer is a general lesson should also be written to the knowledge base, so
the next person searching a symptom finds it. Both, not either. the next person searching a symptom finds it. Both, not either.
- `status: wontfix` is legitimate and requires a sentence saying why. - `status: wontfix` is legitimate and requires a sentence saying why.
+8
View File
@@ -13,6 +13,14 @@ decisions taken over three days; the reasoning is kept, the fragmentation is not
The environment a change is run against before it reaches real machines. The environment a change is run against before it reaches real machines.
> **Still the lab, no longer the test bed — 2026-09-30, by [ADR 0149](0149-the-live-mesh-is-the-test-bed.md).**
> Everything here stands. What changed is what the lab is *for*: a change is verified against the mesh
> that is running, because the faults that cost the most are faults of a mesh that already exists —
> bound consumers, containers made against an older roster, an adopted machine — and a bed is by
> construction a mesh that does not. Raising a mesh from bare is now the lab's whole job, which is the
> one thing the live mesh cannot be asked to do. 0149 also supersedes
> [ADR 0068](0068-the-lab-takes-requests.md), which extended this one and was never built.
## A node in the lab is a virtual machine ## A node in the lab is a virtual machine
It boots a stock Linux image, runs the real install, and becomes a node. **It is not a model of It boots a stock Linux image, runs the real install, and becomes a node. **It is not a model of
+17 -1
View File
@@ -1,6 +1,6 @@
--- ---
topic: building it topic: building it
status: proposed status: accepted
date: 2026-09-01 date: 2026-09-01
deciders: jochen deciders: jochen
reconstructed: false reconstructed: false
@@ -101,3 +101,19 @@ the digest down after building.
**Whether kind 4 deserves a module at all.** Thirty-five descriptions that say *install this and **Whether kind 4 deserves a module at all.** Thirty-five descriptions that say *install this and
write these files* may be better as one module with settings than as thirty-five modules. Left write these files* may be better as one module with settings than as thirty-five modules. Left
open deliberately; it is a question about the shape of the catalogue, not about whether to have one. open deliberately; it is a question about the shape of the catalogue, not about whether to have one.
## Accepted, 2026-09-29, against what was built
*Marked in a grooming pass: the mesh was built to this record and the record still said `proposed`.*
The proposal is the arrangement that exists. `mesh-catalog` holds descriptions of software we did
not write and the programs that provision it, and holds neither the mesh's own components nor an
application's own module. The mesh's list of modules is a table in the control plane, filled by
`module add`, and every module records the source it came from with the commit it was read at.
**One half is not built: `module check` as a command on the control plane's binary.** A manifest is
still validated by a test that reaches into the control plane's internals — which works for this
catalogue and gives nothing at all to somebody describing their own application in their own
repository, which this record says is the case that matters most. That is
[issue 148](../04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md).
@@ -26,6 +26,14 @@ runtime is per-module, not per-node, and treating the audit-logger as special le
modules' code with nothing to run it: the conversion produced tools and events that, as it stands, modules' code with nothing to run it: the conversion produced tools and events that, as it stands,
never execute. never execute.
> **The hosting form is settled elsewhere — 2026-09-30.** Where this record says "a container", read
> [ADR 0150](0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md): a module's own
> code runs as supervised processes under this record's one account. Nothing else here changes — the
> per-module runtime, the per-tool key and the single scoped account are the argument this record made
> and they are why 0150 goes the way it does. The note is here because two design documents chose the
> other form without knowing this record existed
> ([issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md)).
## Decision ## Decision
### A module with tools or events runs a process of its own ### A module with tools or events runs a process of its own
@@ -80,6 +80,15 @@ reaching the routed name, which the clause above has just made resolvable inside
three are one decision: **compose the name, propagate it, certify it** — each is meaningless without three are one decision: **compose the name, propagate it, certify it** — each is meaningless without
the one before it. the one before it.
> **The mechanism changed — 2026-09-30, by [ADR 0148](0148-the-meshs-names-are-resolved-not-copied-into-containers.md).**
> A routed name still reaches every asker in the mesh, which is what this record decided and it stands.
> It no longer reaches them by being written into each declared container: copying the roster in made the
> roster part of every container's identity, so one name moving replaced every container in the mesh
> ([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). A
> container resolves through its machine's resolver instead. The consequence below — that an internal
> issuer's challenge needs the routed name resolvable inside the mesh — holds unchanged, by the means the
> machine itself already uses.
## Consequences ## Consequences
- **Lab-versus-production is one node-level `public-domain` setting**, not an override on every - **Lab-versus-production is one node-level `public-domain` setting**, not an override on every
+8 -1
View File
@@ -1,14 +1,21 @@
--- ---
topic: building it topic: building it
status: proposed status: superseded
date: 2026-09-12 date: 2026-09-12
deciders: jochen deciders: jochen
reconstructed: false reconstructed: false
extends: 0016-the-lab.md extends: 0016-the-lab.md
superseded-by: 02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md
--- ---
# 68. The lab takes requests, one at a time, and runs each from its own copy # 68. The lab takes requests, one at a time, and runs each from its own copy
> **Superseded — 2026-09-30, by [ADR 0149](0149-the-live-mesh-is-the-test-bed.md).** Never built. The
> live mesh became the test bed, because the faults that cost the most are faults of a mesh that
> already exists — bound consumers, containers made against an older roster, an adopted machine — and a
> bed is by construction a mesh that does not. The rule worth keeping from below is that a run reads a
> copy that is not anybody's working tree.
## Context ## Context
**The lab is exclusive hardware, and today a person holds it.** Raising a scenario takes over **The lab is exclusive hardware, and today a person holds it.** Raising a scenario takes over
@@ -35,6 +35,16 @@ The development cycle is enforced mechanically, to the extent frontmatter can ca
capability existed" is an answer). capability existed" is an answer).
- **No silent graduation** — a `graduated` research overview says what it `became:`, and the - **No silent graduation** — a `graduated` research overview says what it `became:`, and the
targets exist. targets exist.
- **No two records answering to one number** — added 2026-09-30; see the insight below.
> **Progressive insight — 2026-09-30.** The list above named four things `cycle.py` enforces, and
> now names five. Nothing enforced that two issue records hold different numbers: two machines
> filing issues within one hour both read `main`, both took "the next free number", and collided
> twice — the second collision reaching `main` with `records.py`, `cycle.py` and `index.py` all
> reporting success (04-ISSUES/155). A number is how every other record cites one, so two records
> answering to it is a citation that resolves to whichever folder the reader opened. `cycle.py`
> refuses it now. The decision here stands exactly as written: this is one more thing frontmatter
> and file names can carry, found by its absence rather than by reasoning.
[`00-META/checks/cycle.py`](../00-META/checks/cycle.py) refuses violations, beside `records.py` [`00-META/checks/cycle.py`](../00-META/checks/cycle.py) refuses violations, beside `records.py`
and `index.py`; all three run before any merge here. What frontmatter cannot see — that code work and `index.py`; all three run before any merge here. What frontmatter cannot see — that code work
@@ -1,6 +1,6 @@
--- ---
topic: what runs on it topic: what runs on it
status: proposed status: accepted
date: 2026-09-25 date: 2026-09-25
deciders: jochen deciders: jochen
reconstructed: false reconstructed: false
@@ -275,3 +275,11 @@ On acceptance, each of these is superseded or amended by this record, not edited
everything a module needs is a requirement everything a module needs is a requirement
- [Issue 095](../04-ISSUES/095-a-module-assigned-after-genesis-has-no-broker-account/00-report.md), - [Issue 095](../04-ISSUES/095-a-module-assigned-after-genesis-has-no-broker-account/00-report.md),
[issue 103](../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md): what fails today [issue 103](../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md): what fails today
## Accepted, 2026-09-29, against what was built
*Marked in a grooming pass.* The vault is a module providing `secret` at mesh scope, and six
modules in the catalogue require it — so a shared secret is a requirement answered by the vault,
which is what this record asks for. Private keys are still made where they are used and never
travel, which is the other half and was never in question.
@@ -1,6 +1,6 @@
--- ---
topic: what runs on it topic: what runs on it
status: proposed status: accepted
date: 2026-09-26 date: 2026-09-26
deciders: jochen deciders: jochen
reconstructed: false reconstructed: false
@@ -189,6 +189,23 @@ On acceptance, each of these is amended by this record, not edited:
credential is no longer all-or-nothing with a window. It overlaps, with each step confirmed. credential is no longer all-or-nothing with a window. It overlaps, with each step confirmed.
A single-party credential is staged, not replaced. A single-party credential is staged, not replaced.
## Accepted, 2026-09-30, and not scheduled
Accepted as written. The separation it draws — a consumer's *resource* and a *credential that reaches
it* are different things with different lifecycles — is the part that had to be settled, because the
alternative is what the record was written against: retiring a credential taking the data it reached
with it. That is a data-loss shape, and a record that names it should not sit unresolved while the
code that could hit it is being written.
**It is not built, and accepting it does not schedule it.** The SDK's provisioner adapter is still
`create` / `remove` / `holds` rather than the four operations above, and no provider implements the
two-credential rotation. Accepted-and-not-built is an ordinary state here — 0141 and 0142 are both in
it — and it is the honest one: leaving this `proposed` made it invisible to anyone reading what the
mesh has decided, while changing nothing about what runs.
The work it implies belongs with the provisioner contract, beside
[issue 124](../04-ISSUES/124-a-consumer-cannot-be-told-what-its-provider-derived/00-report.md).
## Consequences ## Consequences
- **Every credential provider's adapter changes**, in two steps. The first separates *retire a - **Every credential provider's adapter changes**, in two steps. The first separates *retire a
@@ -1,6 +1,6 @@
--- ---
topic: what runs on it topic: what runs on it
status: proposed status: accepted
date: 2026-09-26 date: 2026-09-26
deciders: jochen deciders: jochen
extends: 0112-a-module-definition-names-no-node-mesh-or-path.md extends: 0112-a-module-definition-names-no-node-mesh-or-path.md
@@ -47,3 +47,10 @@ other boundary already is: the module name.
node runs one of each (ADR 0115)" — instead of failing on whichever name collides first. node runs one of each (ADR 0115)" — instead of failing on whichever name collides first.
- Multi-tenant asks are answered in the catalogue (a second module definition), not in the - Multi-tenant asks are answered in the catalogue (a second module definition), not in the
control plane. control plane.
## Accepted, 2026-09-29, against what was built
*Marked in a grooming pass.* The rule is enforced where it cannot be forgotten: `assignment`'s
primary key is `(node, module)`, so a second assignment of one module to one machine is not a thing
the mesh can hold. The record read `proposed` while the schema had already settled it.
@@ -12,7 +12,7 @@ extends: 0102-the-mesh-writes-into-a-shared-file-never-over-it.md
## Context ## Context
When a resource stops being declared — its module unassigned, the node sent a When a resource stops being declared — its module unassigned, the node sent a
deliberately-empty declaration ([issue 127](../04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)), deliberately-empty declaration ([issue 149](../04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)),
or a new catalogue version renaming its id — the host undoes it. The host's own code states or a new catalogue version renaming its id — the host undoes it. The host's own code states
the rule it means to follow: **it removes what it made and leaves what it merely configured.** the rule it means to follow: **it removes what it made and leaves what it merely configured.**
For almost every resource it does exactly that: For almost every resource it does exactly that:
@@ -109,6 +109,23 @@ understated as none. The remaining work is a way to build the host and a way to
path, and until both exist nothing delivers a version and every machine takes the fallback — which is path, and until both exist nothing delivers a version and every machine takes the fallback — which is
what every machine does today. what every machine does today.
> **Progressive insight — 2026-09-30. Both of those exist now.** The paragraph above named two missing
> things and they are built: a Go toolchain, based on a new `mesh-tools-go` module so the compiler is
> named and not pinned, and `${version}` in any value of a resource that uses an archive or a bundle.
> The mesh compiles its own host and publishes it to its own registry, measured — a statically linked
> stripped binary, fetched back out and run. **The cost was larger again than this note said**: three
> more things in the path assumed one language or one shape, and a fourth was in the base image.
> The account is [issue 142](../04-ISSUES/142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md).
>
> The version in a path is the artifact's **digest**, not the commit this note's own wording would
> suggest. Two builds of one commit are meant to be the same bytes, so a content-addressed version
> means an unchanged build keeps the path it had; a commit-named one would move for an identical binary
> and recreate everything reading it.
>
> **Still nothing delivers a version to a machine.** The host is a module and builds, and declares no
> resources, so the bundle sits in the registry and no machine is asked to take it. That is the next
> piece, and the decision above is unchanged by any of this.
## Consequences ## Consequences
- **The host becomes a build target and a module** — a module whose resource is the next host, applied - **The host becomes a build target and a module** — a module whose resource is the next host, applied
@@ -113,6 +113,36 @@ the authority it was bound to, checked in the control plane's own test suite —
that verified anything. That is a weaker thing than the paragraph above describes, and it stays that verified anything. That is a weaker thing than the paragraph above describes, and it stays
written this way until the bed runs. written this way until the bed runs.
> **Progressive insight — 2026-09-30. It has now been run, on the live mesh rather than in the bed.**
> The paragraph above said nothing had verified anything, and something has. The module was registered
> from the catalogue, assigned to a workstation, and checked in the form this section prescribes — the
> authority's own API, so the handshake needs nothing else in the mesh to be right:
>
> ```
> $ curl -sS -o /dev/null -w '%{http_code}' https://<the authority>:9000/health
> 200
> subject=CN=Step Online CA
> issuer=O=Mesh Internal CA, CN=Mesh Internal CA Intermediate CA
> Verify return code: 0 (ok)
> ```
>
> **Both halves.** Unassigning and pushing removed the anchor, emptied the trust store of the mesh's
> authority, and returned the plain client to *unable to get local issuer certificate* — then assigning
> again restored it. The negative half is what distinguishes the anchor working from something else
> having trusted it, and it is the half nothing had ever exercised.
>
> **One thing this found that is not in the module.** The removal only works because the *host* removes
> the service before the script: stopping the unit is what deletes the certificate and refreshes the
> bundles, and it needs the script it calls to still exist. Nothing in the module states that ordering;
> the symmetry this record claims rests on it.
>
> Run on the live mesh because that is where a change is verified now
> ([ADR 0149](0149-the-live-mesh-is-the-test-bed.md)), and the bed still cannot raise a foundation. The
> evidence is [issue 129](../04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/02-resolution.md).
> Extended the same day to every converged machine — `novox`, `g14` and `shanks` each hold the anchor
> and verify with a plain client. `ace` is excluded on purpose: it is adopted, so a module assigned
> there is held rather than run, which is right and is not trust.
## Consequences ## Consequences
The predecessor's authority can be retired from a machine once this module is assigned to it, The predecessor's authority can be retired from a machine once this module is assigned to it,
@@ -0,0 +1,176 @@
---
topic: the tiers
status: accepted
date: 2026-09-30
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0066-public-routing-is-name-agnostic.md
---
# 148. The mesh's names are resolved, not copied into every container
## Context
The mesh gives every container it declares the whole roster of mesh names as entries written into
the container's own hosts file at creation
([design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md),
[ADR 0066](0066-public-routing-is-name-agnostic.md)). A container takes those entries once and never
looks again.
Three issues are the same fact arriving three times.
**A container keeps the address it was made with.** Adopting the predecessor's tunnel moved the hub's
private address; the declaration followed it within one push and nothing on the machine did. The
forge's container held the old address, lost its database, reported healthy while its existing
connections lasted, and then the public name went down
([issue 109](../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md)).
**The same fault, four days later, undetected for five days.** One container had restarted 2286 times
against a database it could no longer find, while the mesh reported the machine as doing what it was
told. Inside it, `novox.internal` was an address that had not existed for five days
([issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md)). Forty-eight
other containers were current, none of them corrected — each had been recreated for some other
reason and picked up the roster on the way.
135 was fixed by putting the roster into the digest the host compares a container against, so a
container whose names moved is recreated like one whose image moved. **That made the roster part of
every container's identity**, which is the third arrival:
**One name moving replaces every container in the mesh.** Migrating one small module on one machine
took four routine actions; each changed the roster, and each replaced every container on the control
node — its own store, the registry, the edge proxy, the forge, the directory, mail. The control plane
was unreachable twice while its own store came back through crash recovery
([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). None of
the replaced containers had anything to do with the module being migrated, or with its machine.
The blast radius of a name is now every container that carries the list, which is all of them. The
node-by-node migration ahead adds names one module at a time — on one machine alone that is around
twenty-five — and each would be a full restart of every service on the hub.
## Considered Options
**1. Keep the roster in every container and accept the churn.** Rejected. It is not a cost that can
be paid down: the mesh gets more names as it grows, and every name costs a restart of everything.
A rollback costs another.
**2. Scope each container's entries to the names it actually binds.** A container is given the names
of the things it declared a requirement on, so a name's blast radius is its consumers. Tidy, needs no
new mechanism, and keeps 135's guarantee exactly.
Rejected, and this is the close one. It contradicts the standing intent that **anything on the mesh
can call anything on it** — three cases, same machine, the private network, the public network, and
no fourth. Scoping resolution to declared couplings makes a name reachable only where the mesh was
told in advance that it would be wanted, and a person debugging inside a container would find names
missing that exist everywhere else on the machine. It also leaves the roster in the digest, so the
churn returns the moment a widely-bound name moves — smaller, not gone.
**3. Resolve at lookup time through the machine's resolver, and copy nothing.** Chosen.
## Decision
**A container resolves the mesh's names through its machine's resolver, at the moment it asks. No
mesh name and no mesh address is written into a container, and none is part of a container's
identity.**
The three consequences that make this worth doing:
- **Staleness stops being possible**, rather than being detected. 109 and 135 are not bugs that were
fixed; they are a shape that no longer exists. A name that moves is answered differently by the next
lookup, in every container, with nothing recreated and nothing restarted.
- **A name's blast radius becomes nothing.** Assigning a module on one machine does not touch a
container on another.
- **Anything can still call anything**, which option 2 gave up. The resolver answers every mesh name to
every asker on the machine, exactly as it answers the machine itself.
**The resolver is a machine-level process, not a container** — one of the modules that is not a
container at all — so a container depending on it is not the circularity it would be if the mesh's
own store had to resolve a name through something the store's own runtime had to start first.
**What a module declares for itself is untouched.** Entries a manifest asks for are the module's own,
stay in the container, and stay in its identity: they are part of what the module *is*, they do not
move when the mesh's roster does, and the mesh does not know what they mean.
**The machine's own roster file is untouched.** It is a file, rewritten in place, read by processes and
people; nothing restarts when it changes. It is only the *copy into each container* that this ends.
### The order this lands in, which is not a preference
**Nothing may stop copying names until resolution works from a container.** Removing the copy first
reintroduces 109 and 135 — silently, and on a live mesh, which is exactly how both were found.
1. **A container on any network can reach the resolver.** Today a container on the runtime's default
network asks from an address the converged filter drops, so it has no DNS at all
([issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md));
and on two of four machines the resolver binds loopback only, so the runtime hands containers a
public resolver instead. Both are prerequisites, not related work.
> **Progressive insight — 2026-09-30, later the same day. The loopback claim was wrong.** The
> resolver bound the private address on all four machines; on two the runtime had never been told
> to use it, and on all four the resolver discarded a query that arrived on the runtime's bridge.
> The step stands; the facts under it were those. Both fixed the same day
> ([issue 110's resolution](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/01-resolution.md)),
> and step 3 landed after them.
2. **The runtime is told which resolver to use, per machine, as a file** — not per container as a
creation-time argument, or the resolver's address is back in every container's identity and the
problem has only got smaller.
3. **Then, and only then, the roster leaves the declaration and the digest.**
Until step 3 the mesh keeps copying, and keeps comparing. 151 stays open until step 3 lands; it is not
closed by this record, only answered by it.
## How this is checked
- **A container resolves a name that moved, without being recreated.** Move a name the mesh serves;
from a container that was running before the move and has not been touched since, the name answers
with the new address. This is the one 109 and 135 would both have failed.
- **A name's blast radius is nothing.** Add a routed name on one machine; no container on any other
machine is recreated. The apply report on each machine says nothing changed. This is 151.
- **Anything calls anything.** From a container on any machine, every `<node>.internal` name and every
routed name the mesh serves resolves — including names the module never declared a requirement on,
which is the guarantee option 2 would have given up.
- **On every network the runtime offers.** The first three hold for a container on the runtime's
default network as well as one on a declared network, because the default network is the case that
has no DNS today.
- **No mesh name is in a container's spec.** A test asserts the digest a host computes for a container
does not move when the mesh's roster does, and does move when the module's own declared entries do.
## Consequences
- **The resolver becomes load-bearing for every container**, where before it was load-bearing for the
machine. This is a real cost and is accepted: a resolver that is down is a machine that cannot
resolve, which is already true of the machine itself, and is a smaller event than a roster change
destroying and recreating every container on the machine.
- **Design 08's "a file rather than a resolver" no longer describes containers.** It was written when
the mesh had no resolver and it gave the right answer then. The reasoning it rested on — every Linux
has a hosts file, no package needed — was already overtaken by names a hosts file cannot express:
service names and wildcards under `<node>.internal`, which is why the resolver was built.
- **ADR 0066's mesh-wide propagation is kept and its mechanism changes.** A routed name still reaches
every asker in the mesh; it reaches them through the resolver rather than by being written into each
container. The consequence 0066 records — that an internal issuer's challenge needs the routed name
resolvable inside the mesh — holds unchanged and by the same means the machine already uses.
- **Issue 110 stops being a container-DNS inconvenience and becomes a prerequisite** for the mesh not
restarting itself whenever it learns a name.
- **A container that names a resolver of its own has opted out of the machine's**, and the copy this
record removes was the only reason such a container could reach anything by a mesh name.
> **Progressive insight — 2026-09-30, the afternoon this landed. Found the hard way.** The mail
> system's admin, behind Mailu's own resolver, lost its database the moment the copy went
> ([issue 171](../04-ISSUES/171-a-modules-own-resolver-knows-no-mesh-name/00-report.md)). A `dns` on
> a container is a decision about whether mesh names exist inside it, not a preference; the module
> was corrected, and whether the controller should refuse the contradiction is that issue's open
> question.
- **A container started by hand gets the mesh's names too**, where before only declared containers did.
Design 08 drew that boundary deliberately, on the grounds that reaching into every container is what
a nameserver would be for. This record accepts that consequence rather than working around it: a
person debugging in a hand-started container resolving the same names as everything else is the
behaviour worth having, and it is what "anything can call anything" means.
## References
- [issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md) — one name replaces every container; the question this answers
- [issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md) — the roster put into the digest
- [issue 109](../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md) — the first arrival
- [issue 110](../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md) — the prerequisite
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — routed names propagate mesh-wide; extended here
- [design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md) — the file-not-resolver reasoning this narrows
@@ -0,0 +1,106 @@
---
topic: building it
status: accepted
date: 2026-09-30
deciders: jochen
reconstructed: false
supersedes: 02-DECISIONS/0068-the-lab-takes-requests.md
---
# 149. The live mesh is the test bed
## Context
[ADR 0068](0068-the-lab-takes-requests.md) proposed that the lab accept queued requests
— a bed and a commit — answer them one at a time from a copy it owns, and expose that through tools so
an agent could start a run and come back to it. It has been `proposed` since 2026-09-12 and nothing was
built.
What happened instead is that the mesh became the thing under test. It runs on four machines; every
fault worth finding in the last month was found on them, and none was found in a bed:
- a container holding an address that had not existed for five days, on the control node
([issue 135](../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md));
- a machine reading healthy for eleven hours while no module could reach another
([issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md));
- one name replacing every container on the hub
([issue 151](../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md));
- a consumer assertion that is correct on a mesh being raised and fatal on one that is running
([issue 156](../04-ISSUES/156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md)).
The last is the one that settles it. That change was exercised on the raise path — which is what a bed
*is* — and the raise path is the only path on which the fault cannot appear. A bed raises a mesh; it
does not have a mesh that has been running for weeks, with consumers already bound, containers created
against an older roster, and an adopted machine carrying a predecessor's configuration. **The faults
that cost the most were all faults of a mesh that already exists**, and a bed is by construction a mesh
that does not.
Lab runs are also expensive in a way that changed the behaviour around them: each costs a build and
several minutes, so they were batched, and a batched test is one whose result arrives after the next
three changes were already written.
## Considered Options
**1. Build 0068 as proposed.** Rejected. It answers a question nobody is asking: the bottleneck was
never that a person had to sit at the lab, it was that a bed cannot hold the state the faults live in.
Queueing and tooling a mechanism that finds the wrong class of fault faster is not an improvement.
**2. Leave 0068 `proposed`.** Rejected, and it is why this record exists rather than nothing. A record
that contradicts current practice and sits unresolved is worse than either answer: it reads as intent
to anyone who finds it, and the practice it contradicts is written down nowhere but a handoff note.
**3. Record that the live mesh is the test bed, and supersede 0068.** Chosen.
## Decision
**A change is verified against the mesh that is running.** Not because a bed would be unwelcome, but
because the state that breaks things is state a bed does not have: containers made against an older
roster, consumers already bound, an adopted machine, a store with weeks of history.
**A change that can only be exercised on the raise path is not verified.** If the only test available
raises a fresh mesh, the record says so, and says which case was therefore not covered. The words
"exercised on a fresh mesh" are a statement about coverage, not a pass.
**The lab is not retired**, and [ADR 0016](0016-the-lab.md) stands. It remains the place to raise a
mesh from bare, which is the one thing the live mesh cannot be asked to do and the one thing a bed does
better than anything else. What this record removes is the lab as the *default* answer to "is this
change good", and with it 0068's queue, tools and request protocol.
**Accuracy over a green run.** An honest failure on the live mesh beats a pass in a bed that could not
have failed — and a change that is risky on the running mesh is a reason to make the change smaller,
not a reason to test it somewhere it cannot break.
**What 0068 got right is kept as a rule, not a mechanism:** a run reads a copy that is not anybody's
working tree. Every run of the lab that mattered was pinned to a checkout rather than a worktree, and
the ones that were not produced results about code nobody had written down.
## How this is checked
- **A record that says a change was verified says on what.** Where it was a fresh mesh, it says which
case is uncovered. This is the clause that would have caught 156: its change was verified, honestly,
on the only path where it works.
- **The lab is not in the path of a merge.** No check, playbook or handoff requires a bed to have run.
- **0068 is unreachable as intent.** Its status is `superseded` and it names this record, so a reader
arriving at the queue design finds out immediately that it was not built and why.
## Consequences
- **A fault can be introduced on the machines that serve.** This is the cost, it is real, and it was
paid twice in one evening — a control plane crash-looping for half an hour, and every container on the
hub recreated five times. Both were found in minutes because they were live, and both would have
passed a bed.
- **There is no pre-merge gate beyond the repositories' own suites.** `make check` and the three hq
checks are what stands between a change and the machines, which raises what those suites are worth
and makes a test that cannot fail a genuine defect rather than an untidiness.
- **Raising a mesh from bare is now the lab's whole job**, and is exercised deliberately rather than
as a side effect of testing something else. The foundation work
([issue 146](../04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md))
is that job, and it is also the proof that the mesh can make another of itself.
- **An agent cannot hand a run to a queue and come back**, which 0068 would have given. In practice it
watches a push and reads the machines, which is what happened anyway.
## References
- [ADR 0068](0068-the-lab-takes-requests.md) — superseded by this
- [ADR 0016](0016-the-lab.md) — the lab, which stands
- [issue 156](../04-ISSUES/156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md) — correct on the raise path, fatal on a running mesh
@@ -0,0 +1,113 @@
---
topic: what runs on it
status: accepted
date: 2026-09-30
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md
---
# 150. A module's own code runs as supervised processes under the module's one account
## Context
The repository answers "what runs a module's own code" two ways and reconciles them nowhere
([issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md)).
[ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) is accepted and
says **a container** — "the tool runtime carrying that module's compiled code" — and "one module, one
process, one account". Two `proposed` design documents say a **`process`** resource running an argv,
supervised by the machine, and one of them declares *four* of them for a single module and presents
four as the point. Neither design document names 0047 in its `decisions:`, and the string `process`
as a resource type appears in no decision record at all. The thing as built is the container.
Two things have happened since 0047 was written that bear on it directly.
[**ADR 0142**](0142-the-mesh-delivers-its-own-components-as-binaries.md) decided that the mesh's own
components are binaries on the machine rather than container images, and
[issue 114](../04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md) was closed
by it. That settled the mesh's components and deliberately said nothing about a module's.
And the standing definition of a module hardened: **a module is software that delivers one or more
services, and a module is not a container.** It may deliver them as a container, an installed package
with a unit, a binary, or configuration files; 61 of 73 happen to use a container and 11 do not,
including the resolver, sshd and fail2ban. A rule that a module's *own code* must be a container makes
the one kind of module the mesh writes itself the only kind that has no choice.
## Considered Options
**1. Hold 0047 as written: a container.** Rejected. Its own reasoning does not require one. What 0047
argued for was a runtime **per module** rather than one for the whole node, because a node-wide runtime
could not hold a per-module broker account and per-module runtimes competing on one tool key would each
be handed calls for tools they do not have. A supervised unit per module satisfies that argument
exactly — it is per module, and a unit runs as an account. The container was the mechanism to hand, not
the conclusion.
**2. Let each design document choose.** Rejected; that is the present state and it is what issue 117
reports. A module author reading the guide writes four processes; a module author reading the record
writes a container; nothing tells either that the other exists.
**3. Settle the hosting form as a supervised process, and settle the count separately.** Chosen.
## Decision
**A module's own code runs as one or more supervised processes on the machine, under the module's single
account.** Where [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) says
"a container, the tool runtime carrying that module's compiled code", read this record. Everything else
0047 decided stands untouched: a tool is served on its own key, only the module that serves it answers,
and the module's account is scoped to exactly its tool keys.
**The invariant is the account, not the process count.** 0047's "one module, one process, one account"
carried its weight in the last clause. Its stated worry about a second process was "not a second one to
scope and seal" — a second *identity* to grant, seal a secret to, and scope on the bus. Several
processes sharing the module's one account create no second identity, so nothing further is scoped or
sealed, and a module may therefore declare as many as its work has shapes: events, tools, a
provisioner, a scheduled ingest. **What a module may not have is two accounts.**
**A module that delivers its service as a container still does.** This record is about the code the
module itself carries — its tools, its events, its provisioner — and not about the software it delivers.
A module wrapping a third-party image wraps a third-party image.
**Why supervised by the machine rather than by the mesh:** it is the same answer ADR 0142 gave for the
mesh's own components, for the same reason. A unit the machine restarts needs no image, no registry
pull and no runtime to be up before the mesh's own code can run — which matters most for exactly the
modules whose code the mesh cannot start any other way.
## How this is checked
- **No design document describes a hosting form for a module's own code without citing this record.**
Designs 18 and 20 name it in `decisions:`; this is the gap issue 117's third point reports, and
`cycle.py` already enforces that a to-be design names its decisions.
- **A module declaring several processes resolves to one account.** A test composes a module with more
than one process resource and asserts the mesh mints exactly one broker account for it, scoped to that
module's tool keys and nothing else — which is 0047's invariant stated as an assertion rather than a
sentence.
- **A module's own code does not require the container runtime.** A machine with no container runtime
can still run a module whose code is its own, which is the claim that separates this from option 1 and
is checkable on a machine that has one by asserting the declaration names no image for it.
## Consequences
- **The sidecar port stops being needed.** [ADR 0029](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md)
records that "anything that is a service plus a sidecar currently has to publish a port to talk to
itself", and the host's `network` shape exists partly for it. A process beside the service on the same
machine reaches it without publishing anything, so that pressure goes.
- **Something must supervise, and it is the machine.** This adds a unit per module's code to what the
host writes and owns. The mesh already writes and owns units — `nftables` proves a module can write one
and run it — so the mechanism exists; the count grows.
- **A module's code is delivered, not pulled**, which puts it behind the same gap as the host's own
delivery ([ADR 0141](0141-the-host-delivers-its-own-successor.md), not built): nothing yet delivers a
version of a module's binary to a machine. A container's code arrives by `docker pull`, and this does
not. **This is the cost of the decision and it is not paid**; until delivery exists, a module whose code
is its own is a module somebody places by hand.
- **Issue 117 is answered and its three disagreements close differently:** container-or-unit is decided
here; one-process-or-several is decided here as several under one account; and whether the record was
consulted is fixed by designs 18 and 20 naming this one.
## References
- [issue 117](../04-ISSUES/117-a-modules-own-code-is-a-container-and-a-process/00-report.md) — the contradiction this answers
- [ADR 0047](0047-a-module-runs-its-code-as-its-own-process-with-its-own-account.md) — extended; its "a container" clause is settled here
- [ADR 0142](0142-the-mesh-delivers-its-own-components-as-binaries.md) — the same answer for the mesh's own components
- [ADR 0141](0141-the-host-delivers-its-own-successor.md) — the delivery this depends on and which is not built
- [ADR 0029](0029-a-network-is-a-shape-because-an-action-cannot-be-undone.md) — the sidecar port this relieves
@@ -0,0 +1,99 @@
---
topic: the tiers
status: accepted
date: 2026-09-30
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0066-public-routing-is-name-agnostic.md
---
# 151. A route's internal name is composed under the node that serves it
## Context
A module that requires a route is given two names from one label: a public one, `<label>.<public
domain>`, and an internal one, `<label>.<node>.internal`
([ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md)). Both were composed
from the node the module runs on.
The two are answered differently. The public name is published into every machine's roster at the
address of the node whose proxy serves it ([ADR 0066](0066-public-routing-is-name-agnostic.md)), so
it reaches the proxy from anywhere in the mesh. The internal name is answered by every machine's
resolver as *anything under a node's name goes to that node*
([design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md)) — the node it was composed from, which is
the consumer's. Where the proxy runs on another machine, that name sends a client to a machine with
nothing listening, while the public name works
([issue 139](../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md)).
Every route on this mesh today is served beside its module, so it has not been seen; `route` is
provided mesh-wide precisely so that stops being true.
Beside it, the roster gave every routed name a second entry with the mesh's suffix appended —
`<name>.<public domain>.internal` — because it composed a full name for every entry as it does for a
machine. That name resolved on every machine, was served by nothing, and was refused by the proxy at
the handshake; the first three names tried while reproducing an unrelated issue were those, and the
evidence pointed at a regression that had not happened
([issue 157](../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md)).
## Considered Options
**1. Keep the consumer's name and publish it at the serving node's address**, as the public name is.
The name stays `<label>.<consumer>.internal` and an exact roster entry overrides the wildcard.
Rejected: it makes `<x>.<node>.internal` mean *goes to that node* except when it does not, which is
the one rule the resolver design states; it needs an entry per route where the wildcard needed none;
and which of an exact entry and a wildcard a resolver answers first is the resolver's business, which
the mesh deliberately does not know.
**2. A proxy on every machine, so the serving node is always the consumer's.** Rejected for this
question: it is a different decision about what `route` is — a node-scoped seat with a mesh-wide
fallback — and this mesh runs one proxy on the hub today. Whatever is decided there, a route served
from another machine must have a name that reaches it.
**3. Compose the internal name under the node that serves the route.** Chosen.
## Decision
**A route's internal name is `<label>.<serving node>.internal` — composed under the node whose proxy
answers the route, which is the machine the request arrives at.** The public name is unchanged:
`<label>.<public domain>` of the node the module runs on, which is where the operator put it.
Where the proxy runs beside the module — every route on this mesh today — the two nodes are one and
nothing changes. Where it does not, the name says where the request goes, which is what a name under
a node's name has always meant.
**A routed name has no mesh form.** The roster publishes it as itself, once, at the serving node's
address. Only a machine has a bare name beside its full one.
What certifies the internal name is unchanged by this: the proxy that terminates it obtains a
certificate from the mesh's authority for the names it is given, and it is given this one.
Taken on the operator's standing instruction to answer the open design questions in the work order.
## How this is checked
- **Composition.** A controller test contributes a route from a module on one node to a proxy offered
from another, gathered the way the controller gathers a consumer's contribution for a provider on
another machine, and asserts the internal name carries the serving node.
- **Publication.** A controller test renders a roster with a machine and a routed name and asserts
the routed name appears as itself, once, and never with the suffix appended.
- **On the mesh.** After the change no machine's roster carries a `<domain>.internal` entry, and a
route's internal name still answers from a container with a certificate from the mesh's authority.
## Consequences
- **A route served from another machine now has a usable internal name.** The first module assigned
that way will resolve, where before it would have resolved to the wrong machine with no error.
- **The internal name of a route can change when its proxy moves.** A route re-homed from one proxy
to another gets a new internal name, as the design's rule implies; clients that dialled the old one
reach the old machine. The public name does not move with the proxy and is the stable one.
- **The roster is one line shorter per routed name**, and a person reading a hosts file no longer
finds names that resolve to a refusal.
- **Issue 139's second question — a per-node route holder — is left open**, and is a decision about
what a seat is rather than about a name.
## References
- [issue 139](../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md) — the question
- [issue 157](../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md) — the alias
- [ADR 0066](0066-public-routing-is-name-agnostic.md) — routed names propagate mesh-wide; extended here
- [ADR 0138](0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md) — how the two names are composed and how far each reaches
- [design 08 §2](../03-DESIGN/01-to-be/08-connectivity.md) — anything under a node's name goes to that node
+34 -5
View File
@@ -48,6 +48,31 @@ form above and dated no earlier than the record's own `date:` — an unmarked ed
violation the reviewer looks for in the diff, and a marked one is legible in the record itself. violation the reviewer looks for in the diff, and a marked one is legible in the record itself.
The git history is the backstop, not the record of intent; the note is the record of intent. The git history is the backstop, not the record of intent; the note is the record of intent.
## A pointer back from what a record changes
A new record naming an old one is not enough. **Where a record changes a mechanism an older record
states — without reversing the decision, so no supersession — the older record gets a dated note
saying where its mechanism now lives.** A reader arrives at the old record by following a citation,
and finds text that is still the decision and no longer the method; nothing in it says a later record
moved the method, and the new record is not in their hands.
> **The mechanism changed — YYYY-MM-DD, by ADR NNNN.** What still stands, what moved,
> and why.
Three examples of the shape, all found by being missed: ADR 0066 still described a routed name being
written into every container after 0148 replaced that with resolution; ADR 0047 still said a module's
code runs in a container after 0150 made it a supervised process; and ADR 0016 still read as though the
lab were the test bed after 0149 said the live mesh is. Each was a citation leading to the wrong
answer, in a record that was not wrong about anything it decided.
**This is not machine-checked, and it cannot be from `extends:` alone.** 102 records extend another and
87 name a parent that does not mention them, which is correct: extending usually means building on a
context, and a one-directional pointer is the right shape for that. What needs a note is the narrower
case where the parent's own text has gone stale, and which case that is, is a judgement — so it is a
rule for the author and the reviewer, and the diff is where it is caught. Making it mechanical would
mean a record declaring the relationship in its frontmatter, which is a change to the record schema and
has not been decided.
The records run in the order the decisions were taken, oldest first. The records run in the order the decisions were taken, oldest first.
**Every decision is a record.** There is no ledger and no index file — if a decision is worth **Every decision is a record.** There is no ledger and no index file — if a decision is worth
@@ -174,6 +199,8 @@ python3 00-META/checks/index.py fail if stale
- **0108** — [A route carries the policy applied to a request, and names a secret rather than holding one](0108-a-route-carries-the-policy-applied-to-a-request.md) - **0108** — [A route carries the policy applied to a request, and names a secret rather than holding one](0108-a-route-carries-the-policy-applied-to-a-request.md)
- **0109** — [A package registry seat is one per ecosystem, not one for all of them](0109-a-package-registry-seat-is-one-per-ecosystem.md) - **0109** — [A package registry seat is one per ecosystem, not one for all of them](0109-a-package-registry-seat-is-one-per-ecosystem.md)
- **0126** — [A module declares its own seats; the mesh reserves its own](0126-a-module-declares-its-own-seats.md) - **0126** — [A module declares its own seats; the mesh reserves its own](0126-a-module-declares-its-own-seats.md)
- **0148** — [The mesh's names are resolved, not copied into every container](0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
- **0151** — [A route's internal name is composed under the node that serves it](0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md)
### What runs on them, and how it gets there ### What runs on them, and how it gets there
@@ -207,9 +234,9 @@ python3 00-META/checks/index.py fail if stale
- **0099** — [A step that runs once names what it reads, and runs again when it changed](0099-a-step-that-runs-once-names-what-it-reads.md) - **0099** — [A step that runs once names what it reads, and runs again when it changed](0099-a-step-that-runs-once-names-what-it-reads.md)
- **0110** — [A seat is held by one assignment, from a closed set, and it may deliver a provision](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) - **0110** — [A seat is held by one assignment, from a closed set, and it may deliver a provision](0110-a-seat-is-a-module-assignment-from-a-closed-set.md)
- **0112** — [A module definition names no node, no mesh and no path: everything it needs is a requirement the mesh resolves](0112-a-module-definition-names-no-node-mesh-or-path.md) - **0112** — [A module definition names no node, no mesh and no path: everything it needs is a requirement the mesh resolves](0112-a-module-definition-names-no-node-mesh-or-path.md)
- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md) *(proposed)* - **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md)
- **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md) *(proposed)* - **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md)
- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md) *(proposed)* - **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md)
- **0117** — [A machine's uplink is a seat: the mesh configures the manager, never the link](0117-a-machines-uplink-is-a-seat.md) - **0117** — [A machine's uplink is a seat: the mesh configures the manager, never the link](0117-a-machines-uplink-is-a-seat.md)
- **0118** — [Undeclaring removes what the mesh made, and gives a unit back the state it was found in](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md) - **0118** — [Undeclaring removes what the mesh made, and gives a unit back the state it was found in](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md)
- **0120** — [A roster fact carries its format as a template: the mesh owns the data, the module owns the format](0120-a-roster-fact-carries-its-format-as-a-template.md) - **0120** — [A roster fact carries its format as a template: the mesh owns the data, the module owns the format](0120-a-roster-fact-carries-its-format-as-a-template.md)
@@ -228,6 +255,7 @@ python3 00-META/checks/index.py fail if stale
- **0145** — [A module checks what the mesh claims is reachable, and it checks itself](0145-a-module-checks-what-the-mesh-claims-is-reachable.md) *(superseded)* - **0145** — [A module checks what the mesh claims is reachable, and it checks itself](0145-a-module-checks-what-the-mesh-claims-is-reachable.md) *(superseded)*
- **0146** — [Connectivity is checked by name, per hosting form, with a valid certificate](0146-connectivity-is-checked-by-name-per-hosting-form.md) - **0146** — [Connectivity is checked by name, per hosting form, with a valid certificate](0146-connectivity-is-checked-by-name-per-hosting-form.md)
- **0147** — [A module anchors the mesh's authority on a machine, and takes it away again](0147-a-module-anchors-the-meshs-authority.md) - **0147** — [A module anchors the mesh's authority on a machine, and takes it away again](0147-a-module-anchors-the-meshs-authority.md)
- **0150** — [A module's own code runs as supervised processes under the module's one account](0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md)
### How it is built ### How it is built
@@ -237,9 +265,9 @@ python3 00-META/checks/index.py fail if stale
- **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md) - **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md)
- **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md) - **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md)
- **0016** — [The lab](0016-the-lab.md) - **0016** — [The lab](0016-the-lab.md)
- **0037** — [Where a module lives](0037-where-a-module-lives.md) *(proposed)* - **0037** — [Where a module lives](0037-where-a-module-lives.md)
- **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md) - **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md)
- **0068** — [The lab takes requests, one at a time, and runs each from its own copy](0068-the-lab-takes-requests.md) *(proposed)* - **0068** — [The lab takes requests, one at a time, and runs each from its own copy](0068-the-lab-takes-requests.md) *(superseded)*
- **0069** — [A module is a repository and a path within it](0069-a-module-is-a-repository-and-a-path.md) - **0069** — [A module is a repository and a path within it](0069-a-module-is-a-repository-and-a-path.md)
- **0076** — [The SDK is a published package, and the toolchain resolves it by version](0076-the-sdk-is-a-published-package.md) - **0076** — [The SDK is a published package, and the toolchain resolves it by version](0076-the-sdk-is-a-published-package.md)
- **0082** — [The registry is reached by name, and the overlay is its security](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md) - **0082** — [The registry is reached by name, and the overlay is its security](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)
@@ -248,6 +276,7 @@ python3 00-META/checks/index.py fail if stale
- **0097** — [A vendor image is a declared build input, and a recipe fetches nothing undeclared](0097-a-vendor-image-is-a-declared-build-input.md) - **0097** — [A vendor image is a declared build input, and a recipe fetches nothing undeclared](0097-a-vendor-image-is-a-declared-build-input.md)
- **0107** — [Persistent data is a directory bind, never a named volume](0107-persistent-data-is-a-directory-bind-never-a-named-volume.md) - **0107** — [Persistent data is a directory bind, never a named volume](0107-persistent-data-is-a-directory-bind-never-a-named-volume.md)
- **0111** — [A build source is on the mesh's git seat, or it is an external repository](0111-a-build-source-is-on-the-git-seat-or-external.md) - **0111** — [A build source is on the mesh's git seat, or it is an external repository](0111-a-build-source-is-on-the-git-seat-or-external.md)
- **0149** — [The live mesh is the test bed](0149-the-live-mesh-is-the-test-bed.md)
### How it is checked ### How it is checked
+39 -2
View File
@@ -7,8 +7,10 @@ code:
- mesh-controller internal/identity/authority.go - mesh-controller internal/identity/authority.go
- mesh-host internal/identity/serving.go - mesh-host internal/identity/serving.go
- mesh-host internal/apply (the service that reflects a rule set) - mesh-host internal/apply (the service that reflects a rule set)
updated: 2026-09-29 updated: 2026-09-30
decisions: decisions:
- 02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md
- 02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md
- 02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md - 02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md
- 02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md - 02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md
- 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md - 02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md
@@ -289,6 +291,31 @@ hosts file by the runtime. That extends the file decision rather than overturnin
mesh and not chosen by a module: a module that listed the machines would go stale the day one mesh and not chosen by a module: a module that listed the machines would go stale the day one
joins, and a module that did not would be one whose containers cannot reach anything by name. joins, and a module that did not would be one whose containers cannot reach anything by name.
**Superseded for containers by [ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
(2026-09-30).** Everything above still describes the machine's own roster file, which is how it works
and how it will keep working. It no longer describes containers.
Copying the roster into each container made the roster part of each container's identity, so one name
moving replaced every container in the mesh — a module assigned on one machine restarted the store, the
registry, the edge and mail on another
([issue 151](../../04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). And it
did not stop the staleness it was meant to: a copy taken at creation is stale the moment the roster
moves, twice found as a container holding an address that had not existed for days
([issues 109](../../04-ISSUES/109-a-container-keeps-the-address-it-was-made-with/00-report.md)
and [135](../../04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md)).
**A container resolves the mesh's names through its machine's resolver, at the moment it asks, and
nothing is copied.** The resolver is a machine-level process rather than a container, so nothing
circular is being asked for. It was gated on a container being able to reach the resolver from any of
the runtime's networks
([issue 110](../../04-ISSUES/110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
and landed the day that did, 2026-09-30: the controller writes no mesh name into a container and the
host's digest carries only what the module declared for itself.
The paragraph below states the old boundary, and 0148 deliberately gives it up: a container somebody
started by hand resolves the same names as everything else, because the resolver answers the machine,
not a list of containers.
**The boundary, which is deliberate and worth stating:** *declared* containers. A container **The boundary, which is deliberate and worth stating:** *declared* containers. A container
somebody starts by hand is not the mesh's to configure, and reaching into every container on a somebody starts by hand is not the mesh's to configure, and reaching into every container on a
machine — declared or not — is what a nameserver in `resolv.conf` would be for. machine — declared or not — is what a nameserver in `resolv.conf` would be for.
@@ -299,6 +326,14 @@ machine — declared or not — is what a nameserver in `resolv.conf` would be f
service, the rest is the node — so what resolves is *anything under a node's name*, going to that service, the rest is the node — so what resolves is *anything under a node's name*, going to that
node. What routes it once it arrives is a proxy's, and stays separate. node. What routes it once it arrives is a proxy's, and stays separate.
*2026-09-30.* **So the node in a route's internal name is the one whose proxy answers it**
([ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md)).
Composed from the node the module ran on, the name sent a client to a machine with nothing listening
whenever the proxy ran elsewhere
([issue 139](../../04-ISSUES/139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md));
composed from the serving node, the rule above holds without exception. The public name stays the
module's node's, which is where the operator put it.
**The mesh writes the data and runs no daemon.** One wildcard per machine, from the same set that **The mesh writes the data and runs no daemon.** One wildcard per machine, from the same set that
writes the hosts file. A resolver is third-party software and runs *on* the mesh rather than being writes the hosts file. A resolver is third-party software and runs *on* the mesh rather than being
*of* it: the mesh has no business shipping one, choosing which one, or knowing its configuration *of* it: the mesh has no business shipping one, choosing which one, or knowing its configuration
@@ -374,7 +409,9 @@ can reach from the outside but cannot resolve from the inside is a name it canno
authority of its own. authority of its own.
**So a granted route is published into internal resolution as well** — the routed name to the node **So a granted route is published into internal resolution as well** — the routed name to the node
that serves it, mesh-wide, by the same mechanism that writes the node names. It is *given by the that serves it, mesh-wide, by the same mechanism that writes the node names — and as itself: a routed
name has no mesh form, and the suffixed alias the roster once added beside it resolved to a refusal
([issue 157](../../04-ISSUES/157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md)). It is *given by the
mesh, not chosen by a module*, for the same reason the node names are: a module listing the routes mesh, not chosen by a module*, for the same reason the node names are: a module listing the routes
would go stale the day one changes. The mesh propagates the names it was told to serve and still would go stale the day one changes. The mesh propagates the names it was told to serve and still
knows nothing about what they mean knows nothing about what they mean
+2 -1
View File
@@ -5,8 +5,9 @@ code:
- mesh-controller cmd/mesh-builder - mesh-controller cmd/mesh-builder
- mesh-controller internal/builder - mesh-controller internal/builder
- mesh-catalog modules/builder - mesh-catalog modules/builder
updated: 2026-09-29 updated: 2026-09-30
decisions: decisions:
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
- 02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md - 02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md
- 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md - 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md
- 02-DECISIONS/0097-a-vendor-image-is-a-declared-build-input.md - 02-DECISIONS/0097-a-vendor-image-is-a-declared-build-input.md
+2 -1
View File
@@ -5,8 +5,9 @@ code:
- mesh-catalog modules/showcase - mesh-catalog modules/showcase
- mesh-controller internal/builder - mesh-controller internal/builder
- mesh-sdk src - mesh-sdk src
updated: 2026-09-21 updated: 2026-09-30
decisions: decisions:
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
- 02-DECISIONS/0099-a-step-that-runs-once-names-what-it-reads.md - 02-DECISIONS/0099-a-step-that-runs-once-names-what-it-reads.md
- 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md - 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md - 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
@@ -155,3 +155,17 @@ design document here, and get it back. That check fails today by design.
**What stands until then** is the signpost, and the honest description of it: reachable, not **What stands until then** is the signpost, and the honest description of it: reachable, not
surfacing. surfacing.
## Where this stands, 2026-09-29
*Added in a grooming pass.* The knowledge base this record is about is the **predecessor's**, and it
is no longer reachable from anything: the surface that answered `recall_search` speaks the transport
the mesh removed at the cut-over
([issue 147](../147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md)).
So the sentence in `README.md` that this record catches — *these documents are still indexed into
the knowledge base* — is now wrong twice over: nothing indexed them, and there is nothing to index
them into. The record stays open, and its answer is no longer "index this repository somewhere"; it
is whatever the mesh grows as its own knowledge surface, if it grows one. Until then the honest fix
is the README, which should stop claiming a property nothing provides.
@@ -1,8 +1,8 @@
--- ---
status: open status: resolved
opened: 2026-09-22 opened: 2026-09-22
located-in: [] located-in: [mesh-controller internal/link/protocol.go (the report field that was missing), internal/inventory, cmd/mesh-controller (node show and status)]
fixed-by: fixed-by: mesh-controller 7683ba8, corrected by 1d9c102
amended-design: amended-design:
--- ---
@@ -0,0 +1,143 @@
# 087 — resolved: the mesh knows which host runs a machine
*2026-09-30.*
## The field existed and was thrown away on arrival
The machine has reported its host version since
[ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) — `Host` on the report, with
a comment saying why it must be there: *"without it nothing can say a machine is behind."*
**The controller's own copy of the report did not have the field.** Two structs describe one message,
one on each side of the wire, and only the sending side had it — so it unmarshalled into nothing and the
mesh could not answer a question the machine had been answering for a week. That is the whole of this
issue's mechanism, and it is worth stating plainly because neither side was wrong on its own.
## What it says now
`node show` names it per machine:
```
last heard from here
host 2026-09-30-0214
```
`not reported — this machine has not said since the mesh began keeping it` where the mesh has not been
told, because a machine that has not said is a different thing from a machine running nothing.
`status` names the machines that are behind another:
```
1 machine(s) run an older host than another machine does:
ace 2026-09-29-0113
the newest any machine reports is 2026-09-30-0214. A host refuses a declaration carrying a
field it does not know, whole — so a new field reaches these machines last
```
## Disagreement, not staleness, and that is deliberate
The open questions asked whether the controller should refuse to send a declaration a node cannot
parse. It cannot yet, honestly: **nothing delivers a host version** (ADR 0141 is accepted and not
built, which is [issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md)),
so the mesh holds no canonical current version and "behind" has no fixed point to be behind.
What it can say truthfully is that these machines do not all run the same host, and which is newest of
the ones it has been told about. That is the fact that matters before a declaration gains a field: **the
oldest host in the mesh is what the mesh may send.**
Two deliberate refusals to guess:
- **A machine that has reported nothing is not called behind.** It may be running anything. `node show`
says it has not said, per machine, which is the honest form.
- **Versions compare as strings.** That suits the timestamps and commits this mesh uses and is wrong
for a scheme where `10` sorts before `9`. Said in the code at the place that would have to learn,
rather than left as a surprise.
## What I shipped first was wrong, and the mesh said so within the hour
The first version reported *"N machine(s) run an older host than another machine does"* and worked out
which by comparing versions as strings. **A host reports its version as a commit, and commits have no
order.**
On the live mesh, with the adopted machine pushed for the first time:
```
3 machine(s) run an older host than another machine does:
g14 04a27ca
novox 04a27ca
shanks 04a27ca
```
Those three run the **newer** host — installed 09:18, against the adopted machine's 01:13. `ced54d4`
sorts above `04a27ca` and that is all it means. An arbitrary lexicographic result, presented as a fact,
about the one thing this record exists to make trustworthy.
The code carried a caveat saying versions compare as strings and that this "is enough for the timestamps
and commits this mesh uses". That was the error, written down and not noticed: it is enough for
timestamps and it is **meaningless** for commits, and the mesh reports commits.
It now reports the split and claims no ordering:
```
4 machine(s) do not all run the same host:
04a27ca g14, novox, shanks
ced54d4 ace
a host refuses a declaration carrying a field it does not know, whole — so the mesh may
send only what every one of these understands. Which of them is newer is not readable
from a commit; that needs a version the host reports as ordered
```
More useful as well as more honest — the reader sees who is on which side of the split, which is what
decides whether a field can be sent — and it leaves the ordering where it belongs: with the host, which
would have to report something ordered for anybody to have it.
**This is [issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
arriving by my own door**, an hour after closing it: a report that confidently says the opposite of the
truth is worse than one that says less, because it trains a reader to distrust the whole surface.
## Measured on the mesh, and one limitation it exposed
All four machines run the identical host binary — same digest, installed within eighteen seconds of each
other — and at first only one reported its version. The other three said *not reported*, which read as a
difference between machines where there was none.
**A machine states its host version only when the mesh sends it a declaration.** The report is published
after an apply; the periodic reconcile that runs every five minutes publishes nothing, because it is the
machine keeping itself as declared rather than answering anything. So a machine that is current and idle
never says, and the mesh cannot distinguish that from a machine running something ancient.
Confirmed by pushing: before, `not reported`; after, `04a27ca` — the same version the machine that had
been pushed already reported.
```
novox host 04a27ca
shanks host 04a27ca
g14 host 04a27ca
ace host not reported — this machine has not said since the mesh began keeping it
```
`ace` has not been pushed since the field existed; it is adopted and parked.
**This is enough for the purpose and not enough for the claim.** For deciding whether a new declaration
field is safe it is sufficient, because pushing is what the mesh is about to do anyway and the answer
arrives with the act. For *knowing what the mesh is running*, it is not: a long-idle machine's entry is
as old as its last push, and the honest reading of `not reported` is "nobody has asked recently" rather
than "this machine is silent". The words say the first, which is why they are those words.
Making a heartbeat carry it would close the gap and is a change to what a heartbeat is — a bare word
that the node is there, deliberately carrying nothing else. Left alone rather than widened in passing.
## The open questions, answered as far as they can be
- *Should a node report the version of its host?* It already did. The gap was the reading.
- *Should the mesh refuse to send a field no node understands yet, or refuse per node and say so?*
Neither, yet — refusing needs the mesh to know which fields need which version, which is the third
question below and is not answered here. What it does is make the disagreement visible before
somebody adds a field.
- *Is there a general shape — a declaration saying which version of the host it needs?* Still open, and
now cheaper to answer: the versions are recorded, so a minimum-version field on a declaration has
something to compare against. It belongs with
[issue 107](../107-a-declaration-carries-no-order/00-report.md), which wants to add a field and is the
first thing this makes safe.
@@ -1,8 +1,8 @@
--- ---
status: open status: resolved
opened: 2026-09-22 opened: 2026-09-22
located-in: [] located-in: [mesh-catalog modules/gitea, mesh-controller internal/catalogue/declaration.go]
fixed-by: fixed-by: mesh-controller 7352c84, merged in #46 — a module is told its port in a container's environment too, as a file already was
amended-design: amended-design:
--- ---
@@ -39,3 +39,9 @@ the assignment happens to differ.
- Should composition refuse an environment value that names a port the module does not fix, the - Should composition refuse an environment value that names a port the module does not fix, the
way it refuses other claims a module cannot make? way it refuses other claims a module cannot make?
- Which other modules write their own address, with a port, into their environment? - Which other modules write their own address, with a port, into their environment?
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* The forge's address follows a moved port the same way every other reader does. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -1,8 +1,8 @@
--- ---
status: open status: resolved
opened: 2026-09-22 opened: 2026-09-22
located-in: [] located-in: [mesh-controller internal/catalogue/declaration.go, mesh-catalog (every routed module)]
fixed-by: fixed-by: mesh-controller bdf965d (a route names the endpoint it serves) with `portOfEndpoint` and `AtPublishedPort` — the contribution carries the endpoint's declared port and the machine-side redirection is applied to it like any other
amended-design: amended-design:
--- ---
@@ -45,3 +45,9 @@ precisely because the predecessor holds the usual one.
keeping the mapping out of rendered configuration? keeping the mapping out of rendered configuration?
- What should refuse a declaration whose contributed route names a port nothing on that node - What should refuse a declaration whose contributed route names a port nothing on that node
listens on? listens on?
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A route names an endpoint rather than a port, and the redirection that turns a declared port into the published one is applied to contributions too. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -1,8 +1,8 @@
--- ---
status: located status: resolved
opened: 2026-09-23 opened: 2026-09-23
located-in: [mesh-controller internal/link, mesh-host internal/link] located-in: [mesh-controller internal/link, mesh-host internal/link]
fixed-by: fixed-by: mesh-host PR 59 (the host refuses an older sequence and drains by it), mesh-controller PR 160 (each send is numbered under the node's hold) — measured 2026-09-30, 02-resolution.md
amended-design: amended-design:
--- ---
@@ -40,3 +40,16 @@ exists there and is thrown away at the wire.
separate genesis-digest branch is needed on the host? separate genesis-digest branch is needed on the host?
- Is a sequence enough, or does a mode change deserve its own marker, so a replayed converged - Is a sequence enough, or does a mode change deserve its own marker, so a replayed converged
declaration is refused by mode as well as by order? declaration is refused by mode as well as by order?
## What has since made this safer to do (2026-09-30)
Adding a `sequence` to a declaration is adding a field, and a host refuses a declaration carrying a
field it does not know — whole. That was
[issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md), and it is resolved: the
mesh now records which host each machine reports and `status` names every machine running an older one
than another does.
So the flag day is visible before it is walked into, which it was not when this was filed. It does not
make the field free: **the oldest host in the mesh is still what the mesh may send**, and one machine of
four is behind today. A sequence that an old host refuses takes that machine out of the mesh's reach
entirely — worse than the replay it prevents, which has never been observed.
@@ -0,0 +1,76 @@
# 107 — diagnosis: the fix is a flag day, and it should wait for delivery
*2026-09-30. Read, measured, and not built — deliberately.*
## The premise is confirmed
A host parses a declaration with unknown fields refused, and the code says why rather than leaving it
to be inferred:
> `DisallowUnknownFields` is the whole point rather than strictness for its own sake: a field the host
> does not know is a thing the control plane believes it asked for.
So adding `sequence` and `supersedes` is not an additive change. **Any host that has not been upgraded
refuses the whole declaration and applies nothing** — which is exactly the behaviour that keeps a
half-understood declaration off a machine, and exactly what makes this expensive.
## What has changed since this was filed
[Issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md) is resolved: the mesh now
records the host version each machine reports and `status` names every machine running an older host
than another does. The flag day is visible before it is walked into, which it was not on 2026-09-23.
That makes the cost measurable rather than hypothetical, and the measurement is the reason this is not
being built today.
## Why it waits
**One machine of four runs an older host, and it cannot be upgraded.** `ace` is adopted, deliberately
parked until the network and module-assignment work is settled, and **nothing delivers a host version at
all** — [ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) is accepted and not
built, which is [issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md).
Every machine takes a hand-placed binary.
So shipping the field means, in order: place a host by hand on three machines, unpark the fourth, place
it there too, and only then turn the controller half on. A machine missed in that sequence is a machine
the mesh cannot send anything to at all — not degraded, unreachable.
**And the fault it prevents has never been observed.** The record says so itself: *"Not observed;
constructed from the code, and narrow."* It needs a backlog of more than sixteen declarations queued
across a `converge`/`adopt` pair, or a broker slow enough to split one, and the host already applies the
newest of a drained batch and refuses a declaration that is not the last by digest.
**Trading a machine's reachability for a replay nobody has seen is the wrong way round.** The right
order is [issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md) first —
when the mesh can deliver a host, a declaration field costs a rollout instead of an expedition — and the
work order already puts that in its last group, as the proof that the mesh can make another of itself.
## What the open questions look like now
- *A per-node `sequence` under the controller's node hold, and `supersedes` as the previous digest?*
Still the right shape. The controller already holds the lock and already records each send, so the
order exists and is thrown away at the wire — unchanged since this was filed.
- *Genesis signing its bundle as sequence zero?* Yes, and it is the cheaper half: the bundle is written
by the host that will read it, so it has no flag day of its own.
- *Is a sequence enough, or does a mode change deserve its own marker?* A sequence alone does not stop
a replayed *converged* declaration reaching a node that has since been returned to adopted, which is
the incident of issue 104 by another door and is what this record names as its real risk. It wants
both, and the second is the one worth having first.
- **And one this record did not ask:** should a declaration say which host version it needs? 087 makes
that comparable for the first time, and it is the general form of the answer — a field that announces
its own requirement, rather than a flag day per field, for ever.
## Status
Left `located`. The owner is unchanged, the shape of the fix is agreed, and the gate is
[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md) rather than anything
in this record. **This is a judgement about order, not a refusal** — it is cheap to overrule, and the
code is a day's work once a host can be delivered.
## The gate has opened (2026-09-30, evening)
The mesh delivers the host now — built by its own toolchain, published to its own registry, delivered
over the bus and started by the launcher, on all four machines
([issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md)). A declaration
field is a build and a push, not an expedition. The order this record asked for — hosts first, then
the controller — is now two commands and a status line that says when the first has finished.
@@ -0,0 +1,60 @@
# 107 — resolved: a declaration carries its order
*2026-09-30. Measured on the mesh.*
## What was done
**Hosts first, then the controller** — the order [issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md)
says a new declaration field needs, and now a build and a push rather than an expedition
([issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/01-progress.md)).
The host understands a `sequence` on a declaration and tolerates its absence: absent reads as "no
order claimed", not "first", so a controller that sends none is still understood and a host that kept
a declaration before it understood the field compares nothing. It refuses a declaration with a lower
sequence than the one it kept, whole, and says why; and the drain that picks one declaration from a
batch keeps the highest sequence rather than the last to arrive — which is the case the report
constructed, a backlog drained out of order.
The controller numbers each send: the next number for that node, taken under the node's hold, before
the body exists, so the number is inside what the mesh signs and a replayed older declaration cannot
borrow a newer one's.
## Measured
```
push shanks; push shanks
sequence in kept declaration: 2
node sequence
novox 2
shanks 2
ace (none — not sent since numbering)
g14 (none)
status: nobody "not running what the mesh would send them"
```
Both applies went through; neither was refused; the machine holding the earlier one accepted the later.
## The subtlety, which would have read every machine as behind for ever
The mesh decides a machine is behind by comparing the digest of what it **would** send against what it
**did** send. A number changes the bytes. So the read-only comparison composes with the number the
machine was *last* sent — not a fresh one — and is byte for byte what was sent when nothing else
changed. Without that, numbering would have made `status` name all four machines as out of date on
every reading, permanently.
## The open questions
- *A per-node `sequence` under the controller's node hold?* Yes, as described. **`supersedes` — the
previous digest — is not added.** A strictly-greater sequence gives the ordering; a chain of digests
would give continuity, which nothing here needs yet and which every re-composition would break.
- *Genesis signing its bundle as sequence zero?* Zero is "no order claimed", which is what the bundle
carries by carrying nothing. Same rule, no genesis branch.
- *A marker for a mode change?* Not needed for the incident it guards: a replayed converged declaration
reaching a node returned to adopted is already refused **by mode**, before this check runs.
## How it is checked
Host: an older sequence is refused, a newer or equal one is not, and no order claimed on either side
compares nothing; the drain keeps the highest sequence, and falls back to arrival when none is claimed.
Controller: a send carries its number inside the signed bytes, an unnumbered send is byte for byte what
it was before, and each node's counter is one higher per send and readable for the comparison.
@@ -1,8 +1,8 @@
--- ---
status: located status: resolved
opened: 2026-09-24 opened: 2026-09-24
located-in: [mesh-host internal/apply] located-in: [mesh-controller internal/catalogue/declaration.go (every container was given the roster at creation)]
fixed-by: fixed-by: mesh-controller PR 161 — no container is given a mesh name; it resolves through its machine's resolver (ADR 0148, landed 2026-09-30 once issue 110 did)
amended-design: amended-design:
--- ---
@@ -59,3 +59,27 @@ container is made with are the same kind of input, read once at creation, and ar
and is the stronger statement; it is also what the mesh's own resolver exists for. and is the stronger statement; it is also what the mesh's own resolver exists for.
- Either way: what tells an operator that a container is running with an address the node no longer - Either way: what tells an operator that a container is running with an address the node no longer
has? Nothing did. has? Nothing did.
## Answered at the cause (2026-09-30)
This was the first of three arrivals of one fact: a container is given the mesh's names when it is
created and never looks again, so a name that moves afterwards is wrong inside it for as long as it
runs. It arrived again as [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md),
whose fix made the names comparable — and that fix made the roster part of every container's identity,
which arrived as [issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md).
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) ends the
copying: a container resolves through its machine's resolver at the moment it asks. The shape this
record reports then has nowhere to occur. It is gated on
[issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md), so
until that lands the mesh still copies and still compares.
## Resolved (2026-09-30)
110 landed the same day ([its resolution](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/01-resolution.md)),
and mesh-controller PR 161 then removed the copy: no container is given a mesh name or a mesh address,
and a module's own declared entries are the only `host` lines it carries. Verified on the control-node
after its containers were recreated once — the last time a name will do that: the forge's container
carries no extra hosts and resolves another machine and a routed name through the machine's resolver,
so the shape this record describes has nowhere to occur. Checked in the controller's tests: a
container's declaration is byte-for-byte the same under a roster of one machine and a roster of three.
@@ -1,8 +1,9 @@
--- ---
status: open status: resolved
opened: 2026-09-24 opened: 2026-09-24
located-in: [] located-in:
fixed-by: - mesh-catalog modules/dnsmasq (the runtime was never told; the resolver answered by interface)
fixed-by: mesh-catalog PR 175 (the runtime is reloaded and keeps its containers over a restart) and PR 176 (the resolver answers by address, so a query from a bridge is admitted) — measured 2026-09-30, 01-resolution.md
amended-design: amended-design:
--- ---
@@ -51,3 +52,22 @@ knows that is what the rule means.
the runtime's default one? That is a stronger rule and would have prevented 109 as well. the runtime's default one? That is a stronger rule and would have prevented 109 as well.
- What checks it? A converged bed with a container on the default network resolving a mesh name is - What checks it? A converged bed with a container on the default network resolving a mesh name is
the missing assertion; nothing in the resolver's own beds covers the filter. the missing assertion; nothing in the resolver's own beds covers the filter.
## What now depends on this (2026-09-30)
This stopped being a container-DNS inconvenience.
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) decides
that a container resolves the mesh's names rather than being given a copy of them, which is what stops
one name moving from replacing every container in the mesh
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)) and what makes a
stale address impossible rather than merely noticed
([issues 109](../109-a-container-keeps-the-address-it-was-made-with/00-report.md)
and [135](../135-a-containers-mesh-names-are-not-compared/00-report.md)).
**That decision cannot land until this one does**, and not partly: a container on the runtime's default
network is the case with no DNS at all, and it is the case the mesh's own forge runs in. Two of four
machines also bind the resolver to loopback only, so the runtime hands their containers a public
resolver. Both halves are this issue.
*Later the same day: the second half was wrong, and the first had a different cause than the one above.
[01-resolution.md](01-resolution.md) has what was actually found.*
@@ -0,0 +1,64 @@
# 110 — resolved: a container on any network reaches the resolver, and is answered
*2026-09-30. Measured on the three converged machines; the adopted one holds its resolver module until it
is taken and is not covered.*
## What was actually wrong
Not what the report predicted. The report named the filter: a container on the runtime's default
network asks from a bridge address, and the converged filter admitted queries by source address only.
That was true when it was written and was fixed before this issue was ever tested — the filter admits
by the link a packet arrives on ([ADR 0144](../../02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md)),
and a container's bridge is admitted whole. Tested on every machine: the query arrives, the filter
passes it.
Three other things were wrong, each hiding the next.
**The runtime had never been told.** The resolver module writes the runtime's `dns` key into the
runtime's own configuration file. The runtime reads that key when it starts and not on a reload, and on
two machines the runtime predated the file — so every container they started got a public resolver, and
`novox.internal` came back as not existing. Nothing reported this: the file was present and current,
the resolver ran, and a name not existing is a valid answer. Fixed in mesh-catalog PR 175: the module
also sets `live-restore` and reloads the runtime when its file changes, so the one restart the `dns` key
needs no longer stops every container. The restart is then the operator's, once per machine; done on
both today, with every running container kept.
**The resolver dropped the query.** With the runtime corrected, a container's query reached the resolver
— and got no answer, on every machine, including the one whose runtime had been right all along. The
socket was bound to the private address; the filter admitted the packet; dnsmasq received it and
discarded it without a line of log. Its configuration said `interface=mesh0`, and dnsmasq admits a
query by the interface it arrives on when told an interface: a container's query is addressed to the
private address but arrives on the runtime's bridge, and the bridge is not `mesh0`. Fixed in mesh-catalog
PR 176: the resolver is told the address to answer on, not the interface that carries it, and a query to
that address is admitted whatever bridge brings it. The bridges are the runtime's to name.
**The report's second half was wrong.** "Two of four machines bind the resolver to loopback only" was
an inference from the containers' behaviour, and the behaviour had the cause above. The resolver bound
the private address on all four; nothing had asked it there.
## What is verified
From a container on the runtime's default network, started by hand and given nothing, on each of the
three converged machines: `novox.internal` answers with the hub's private address, through the machine's
own resolver. That is the fourth check of
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) — "on every
network the runtime offers" — and its first step; the record's step 2 (the runtime told per machine, as
a file) was already how the module works. Step 3 may now begin.
## What checks it
By hand, today. Nothing in the mesh asserts that a container can resolve a mesh name: the resolver's
own tests cover what it answers, not who can ask. The check that would have caught all three faults is
the one the report asked for and 0148 lists — a container on the default network resolving a mesh name
— and it is not built. It belongs with the reachability check of
[issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
which is parked; until then this is a thing a person verifies after touching the resolver, the filter,
or the runtime's configuration.
## What this cost to find
The three faults produced one symptom — a container that cannot resolve — and each fix revealed the
next. The first was found by reading the runtime's own view of its configuration rather than the file;
the second by capturing the query on the bridge and finding it arrive and go unanswered; the third only
by admitting the first belief was wrong. A machine that had been believed to work all day had never
worked either.
@@ -1,8 +1,8 @@
--- ---
status: open status: resolved
opened: 2026-09-24 opened: 2026-09-24
located-in: [mesh-controller module.json, mesh-host internal/apply] located-in: [mesh-controller module.json, mesh-host internal/apply]
fixed-by: fixed-by: ADR 0142 — the mesh's own components are delivered as binaries on the machine, so the controller is a process; the delivery itself is issue 142
amended-design: amended-design:
--- ---
@@ -84,3 +84,21 @@ restart and run-to-completion semantics — so this would not need host-side wor
The two do not collapse into one. The controller is not a code-carrying sidecar, and `network: host` The two do not collapse into one. The controller is not a code-carrying sidecar, and `network: host`
is what makes the asymmetry visible here and nowhere else. is what makes the asymmetry visible here and nowhere else.
## Answered
*2026-09-29, in a grooming pass.* This asked a question rather than reporting a defect, and the
question was taken: [ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md)
decides that the mesh's own components — the host, the controller, the catalogue, the builder, the
vault — are **binaries on the machine**, delivered by the mechanism
[ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) built, and that
third-party software (the store, the registry, the broker) stays a container because an image is the
right way to carry somebody else's build.
So the operating experience this record was written from — every mutating command reached through
`docker exec mesh-controller` — is answered, and answered against the container.
**The delivery is a separate matter and is not this record's.** Step 1 of it is built and no
component travels yet; that is
[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md).
@@ -1,8 +1,8 @@
--- ---
status: located status: resolved
opened: 2026-09-25 opened: 2026-09-25
located-in: [hq, mesh-catalog modules/showcase, mesh-sdk src/tools/index.ts, mesh-tools] located-in: [hq, mesh-catalog modules/showcase, mesh-sdk src/tools/index.ts, mesh-tools]
fixed-by: fixed-by: hq 83791f0 (PR 196) — ADR 0150: a module's own code runs as supervised processes under the module's one account; designs 18 and 20 now cite it, and ADR 0047 carries a dated note pointing at it
amended-design: amended-design:
--- ---
@@ -99,3 +99,27 @@ what the sidecar is has nowhere in the design layer to look, which is how this w
is mechanically checkable: the resource types a design doc names are a closed set, and every is mechanically checkable: the resource types a design doc names are a closed set, and every
member of it either appears in a decision or does not. Whether that check is worth writing is member of it either appears in a decision or does not. Whether that check is worth writing is
part of this issue, not settled by it. part of this issue, not settled by it.
## Answered (2026-09-30)
[ADR 0150](../../02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md)
settles all three disagreements, and the design documents win two of them:
1. **Container or unit — a supervised process.** ADR 0047's argument never required a container. It
argued for a runtime *per module*, because a node-wide one could not hold a per-module account and
per-module runtimes on one tool key would be handed calls for tools they do not have. A unit per
module satisfies that exactly, and a unit runs as an account. The container was the mechanism to
hand. [ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md) had
already gone the same way for the mesh's own components, and a module is not a container.
2. **One process or several — several, under one account.** 0047's "one module, one process, one
account" carried its weight in the last clause; its stated worry was "not a second one to scope and
seal", which is about a second *identity*. Processes sharing the module's one account create none.
What a module may not have is two accounts.
3. **Whether the record was consulted — fixed rather than answered.** Designs 18 and 20 now name 0150
in `decisions:`, and 0047 carries a dated note saying where its hosting form was settled, so neither
door leads to the wrong answer any more.
**What this does not fix.** A module's code delivered as a binary is behind the same gap as the host's
own ([ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md), accepted and not
built): a container's code arrives by `docker pull` and this does not, so until delivery exists such a
module is one somebody places by hand. 0150 records that as the cost of the decision, unpaid.
@@ -1,8 +1,8 @@
--- ---
status: located status: resolved
opened: 2026-09-25 opened: 2026-09-25
located-in: [mesh-catalog modules/umami] located-in: [mesh-host internal/apply (the spec comparison, via issue 135) — not mesh-catalog modules/umami, which this record first named and which was never at fault]
fixed-by: fixed-by: mesh-host e82789a (issue 135's fix) — recreating the container with a current roster ended it; verified live 2026-09-30
amended-design: amended-design:
--- ---
@@ -0,0 +1,79 @@
# 118 — resolved: it was issue 135, and it is over
*Verified on the machine, 2026-09-30.*
## It no longer happens
```
$ docker inspect -f '{{.RestartCount}}' umami
0
$ docker logs umami --tail 12
26 migrations found in prisma/migrations
No pending migrations to apply.
✓ Database is up to date.
▲ Next.js 16.3.4
✓ Ready in 0ms
$ curl -o /dev/null -w '%{http_code}' https://umami.novox.be/
200
```
No restart loop, no `prisma.$queryRaw()` timeout, and the public name that had answered `502` since
2026-09-25 answers `200`. The raw query that could not complete now runs twenty-six migrations and
reports the store up to date.
## What it was
**The same fault as [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md), which
was diagnosed three days later without either record noticing the other.** 135's container is this
one, named in its own evidence table:
```
umami created 2026-09-23 novox.internal:10.42.0.1
mesh-catalog created today novox.internal:10.10.0.1
```
The mesh's overlay range had moved. Umami had been created before the move and held the store's name
at an address that no longer existed, while every container made after the move held the current one.
That is why the dial appeared to succeed and the first real query timed out, and why the same query
from the same network with the same credential answered in milliseconds — **what differed was the
name, not the path, not the credential and not the store.**
Today it holds `novox.internal:10.10.0.1`.
## Why this record did not find it
The report's own reasoning is worth keeping as a warning, because it is careful and it is wrong:
> A dial that succeeds and a query that times out, from a container on one network to a store on
> another, has the shape of a path-MTU or conntrack fault (large response packets dropped after the
> small handshake ones pass), or of the store accepting the TCP connection while the backend it
> proxies for is wedged.
Both are good hypotheses about a network path. Neither is the answer, and the record also names the
move that would have found it — *"what is known to differ for umami against every working consumer of
the same store tonight is nothing yet — that comparison is the first move"* — and then did not make it.
Comparing umami's hosts entries against any container created that week would have shown a five-day-old
address in one field.
**A stale name presents as a network fault.** That is the lesson, and it is the reason
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) stops
copying names into containers at all: not because detecting staleness is hard, but because it disguises
itself as something else for five days while every check reports success.
## What actually ended it
Issue 135's fix — `mesh-host` e82789a, *a container's mesh names are part of what it is* — put the
roster into the spec digest the host compares, so a container whose names moved is recreated like one
whose image moved. That recreated umami with a current roster and ended this.
That fix has since been superseded in turn, by 0148, because making the roster part of every
container's identity meant one name moving replaced every container in the mesh
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). So this record
closes on a fix that is itself on the way out — which does not make it less closed, and is worth saying
plainly rather than leaving a reader to find out.
## Not carried forward
The `502` had one other contributor worth recording as ruled out: a stale duplicate Traefik router for
this name, hand-authored in the predecessor's era beside the mesh-written one, removed on 2026-09-25
and — as the report says — changing nothing. It was not the cause and it is gone.
@@ -1,8 +1,8 @@
--- ---
status: located status: resolved
opened: 2026-09-26 opened: 2026-09-26
located-in: [mesh-sdk src/provisioner, mesh-catalog modules/redis] located-in: [mesh-sdk src/provisioner, mesh-catalog modules/redis]
fixed-by: fixed-by: mesh-catalog bbda88c, merged in #84 — every credential provider says whether it still holds a consumer
amended-design: amended-design:
--- ---
@@ -60,3 +60,9 @@ checks it after the first pass.
instance and leaves the gap for the others. instance and leaves the gap for the others.
- Where does the record of what was applied live, if not in memory? ADR 0114, still - Where does the record of what was applied live, if not in memory? ADR 0114, still
proposed, puts rotation state with the vault. The same place may answer this. proposed, puts rotation state with the vault. The same place may answer this.
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A provisioner asks the backend what is there rather than trusting what it remembers doing. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -1,7 +1,9 @@
--- ---
status: open status: resolved
opened: 2026-09-26 opened: 2026-09-26
located-in: [mesh-host internal/apply, mesh-controller] located-in: [mesh-host internal/link (the apply report line), mesh-controller cmd/mesh-controller (status and its all-well condition)]
fixed-by: mesh-host cbf5018, mesh-controller bfd983e
amended-design:
--- ---
# 125 — a hold is not a line in the apply report, and an operator flew blind into an outage # 125 — a hold is not a line in the apply report, and an operator flew blind into an outage
@@ -0,0 +1,70 @@
# 125 — resolved: a hold is a line in the report, and it stops the mesh reading as well
*2026-09-30.*
## What the four surfaces say now
The report named four surfaces, none of which carried the one sentence that mattered. Two of them
already did by the time this was picked up, and two did not.
**1. The apply report says what it held** — this was missing, and is the line the operator was reading
when the count did not add up:
```
applied 330 resource(s), 16 held until their module is taken (route-proxy: 13, mailu: 3)
```
Grouped by module and ordered by name, because `take` acts on a module and that is the sentence an
operator needs. An apply that held nothing says nothing extra — a line reporting `0 held` on every
converged apply is one that stops being read.
**2. `status` counts holds, and a hold breaks "all well"** — this was missing. Status now says:
```
15 resource(s) are held as found, because their module was assigned and never taken — so it is
running none of what it declares:
novox route-proxy (13), mailu (2)
`take <node> <module>` compares what runs against what it declares, and runs it
```
**And the sentence that was the fault no longer prints.** "all doing what they were told, all heard
from, running what the mesh would send them" was *true* for the whole outage, and acting on it stopped
the predecessor's proxy. A hold now suppresses it; being adopted still does not, and the difference is
deliberate — adopted is a mode somebody chose and can leave alone, a module assigned and never taken is
a half-finished action with nothing left to finish it.
**3. `node show <node>` shows the node's own held list** — already true, and recorded here as checked
rather than assumed. It reads `held` from what the machine last reported, with an `as of` beside it, and
names each hold's kind, target, id, module, whether something other than the mesh has changed it, and
where an original was kept.
**4. The push's count** is unchanged and now interpretable, which was the ask: `sent 346` against
`applied 330, 16 held until their module is taken (…)` is a pair a reader can resolve without opening a
file on the machine.
## Where the data came from
**The host already reported it.** `Held` has been on the wire since [ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md),
and the controller already recorded it and showed it in `node show`. Nothing needed a new field, a new
message or a migration — which is why this is additive, and why the report's framing (*"the semantics
are consistent and right; the reporting is what let them be forgotten"*) was exactly right.
What was missing was that two surfaces never asked. Status read the mesh's take-time listing, so a
module assigned after that listing showed nothing at all; the apply line counted what it applied and
said nothing about the difference.
## One thing deliberately not done
**Status does not call a hold a fault.** It is correct behaviour, and a reader trained to see red for
something the mesh did right will stop reading. It is reported as work outstanding, with the command
that finishes it — and it withholds the all-well sentence, which is the part that carries the weight.
## How it is checked
- The apply line names the count and the module, and is empty when nothing is held (mesh-host).
- Status finds a hold from what the machine reported, end to end through the store.
- **A held module makes the all-well condition false**, asserted against the production condition
rather than a copy of it — that condition is now one named function for this reason.
- The JSON form carries a row per machine and module, and omits the field entirely when nothing is
held.
@@ -1,5 +1,6 @@
--- ---
status: located status: resolved
fixed-by: mesh-host 1cb8953 and fdc768c — the mesh writes into a marked block of a text file instead of over it
opened: 2026-09-26 opened: 2026-09-26
located-in: [mesh-controller internal/catalogue/facts.go, mesh-host internal/apply] located-in: [mesh-controller internal/catalogue/facts.go, mesh-host internal/apply]
--- ---
@@ -71,3 +72,9 @@ private network loses that name too.
- The host's file resource supports `into: "json"` only; anything else is a whole write. - The host's file resource supports `into: "json"` only; anything else is a whole write.
- `node show <node>` on the adopted workstation: `holds file /etc/hosts - `node show <node>` on the adopted workstation: `holds file /etc/hosts
mesh-wireguard.fact-node-names`, original kept. mesh-wireguard.fact-node-names`, original kept.
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* A shared hosts file keeps every line that is not the mesh's. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -1,7 +1,8 @@
--- ---
status: located status: resolved
opened: 2026-09-26 opened: 2026-09-26
located-in: [mesh-catalog ca-trust] located-in: [mesh-catalog modules/ca-trust]
fixed-by: the ca-trust module registered from the catalogue and assigned to a workstation; verified in both directions 2026-09-30 (02-resolution.md)
--- ---
# 129 — nothing makes a machine trust the mesh's own certificate authority # 129 — nothing makes a machine trust the mesh's own certificate authority
@@ -1,32 +1,82 @@
# Diagnosis # 129 — diagnosis
*2026-09-29.* *2026-09-30, from the workstation the issue was opened on.*
## What was ruled out ## Still live, and reproduced exactly
**That something already carries the root and it is only misplaced.** It does not. The authority The certificate is genuine, the authority is the mesh's, and nothing on the machine trusts it:
serves its root at a path beside its ACME directory, and the one thing that fetches it — the route
proxy — puts it in a directory of its own and hands it to one program. Nothing has ever written
into a machine's trust store. Measured on three converged machines: the anchors present are the
predecessor's authority and a developer tool's local root, and on the machines where the
predecessor's was deliberately removed, every internal name fails verification.
**That the private network could carry it, the way it carries the registry's trust.** That is what ```
the report proposed, and it was rejected on consideration rather than on difficulty $ openssl s_client -connect keycloak.novox.internal:443 -servername keycloak.novox.internal
([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md), option 1): being on subject=CN=keycloak.novox.internal
the network is what makes the registry *reachable* and is therefore the right trigger there, while issuer=O=Mesh Internal CA, CN=Mesh Internal CA Intermediate CA
trusting an authority is a separate fact from being able to reach it. The anchor's directory and Verify return code: 20 (unable to get local issuer certificate)
the command that refreshes the extracted bundles are also one operating system's difference, which
is the host's half of the mesh and not the controller's.
**That it needs a new host resource type.** It does not, today. A file and a service say the whole $ curl https://keycloak.novox.internal/
of it, which the packet filter already proves. The primitive becomes the right answer when a second curl: (60) SSL certificate OpenSSL verify result: unable to get local issuer certificate (20)
operating system is in play, and not before. ```
## Where it belongs `trust list` holds no entry for the mesh. The anchors present are two `mkcert` development roots and
the predecessor's lab root — the report's account of the trust store is unchanged.
A module in the catalogue: it requires `internal-acme-ca`, fetches the root over the mesh's own The public name on the same proxy verifies cleanly (`CN=keycloak.novox.be`, Let's Encrypt, return code
network, installs it as a trust anchor, refreshes the machine's bundles, and — because being 0), which places the fault exactly where the report puts it: not in the proxy, not in the authority,
unassigned stops its unit, and stopping the unit is what undoes it — takes both away again. and not in the certificate.
The owner is therefore `mesh-catalog`, module `ca-trust`, and nothing in the control plane. **The name matters, and the report's "every HTTPS name the mesh serves internally" is too broad.** The
served internal name is `<label>.<node>.internal`. The hosts file also carries
`<label>.<public-domain>.internal`, which nothing serves and which fails differently — that is
[issue 157](../157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md), found while
reproducing this, and it cost the first several minutes of this diagnosis.
## The authority serves what the module needs
`step-ca` is up and healthy, and publishes `roots: /roots.pem` for both `acme-ca` and
`internal-acme-ca`. That endpoint returns PEM:
```
$ curl -sk https://127.0.0.1:9000/roots.pem
-----BEGIN CERTIFICATE-----
MIIBvzCCAWWgAwIBAgIQYa2CkdJk16JyG/dVy2qoEzAKBggqhkjOPQQDAjA+…
```
So `${bound:internal-acme-ca:roots}` in the `ca-trust` module composes to a URL that returns a
certificate, and the module's own check — refuse a body that is not one — is checking the right thing.
Worth recording because it was nearly filed as a defect: step-ca *also* serves `/roots`, which returns
`{"crts":["-----BEGIN CERTIFICATE-----\n…"]}`. That body contains the literal text the module greps
for, so had the module used `/roots` it would have installed JSON into the anchors directory and
reported success. It does not use it. The guard is sound only because the published path is the PEM
one, which is worth knowing before anybody changes either.
## What is actually in the way
**The module is not registered.** The report and the work plan both say it exists and is merged, which
it does — `mesh-catalog modules/ca-trust`, on `main`. But the mesh has never been told about it:
```
$ mesh-controller module list | grep -iE 'ca-trust|step-ca'
step-ca 1 built 67f5f4cf on novox
```
39 of the catalogue's 76 manifests are registered. `ca-trust` is one of the 37 that are not, so it
cannot be assigned to anything — "assign it to one machine" has no module to name.
A dry run confirms it registers cleanly and needs no artifact built: it declares a directory, a script,
a unit and a service, and no image.
```
$ mesh-controller build <catalogue> --path modules/ca-trust --dry-run
… the manifest, parsed and validated
```
## So the remaining work is three steps, not one
1. **Register it** — build it from the catalogue, which pins nothing because it has no artifacts.
2. **Assign it** to a machine. The workstation this was observed on is the honest first choice: it is
where a person meets the fault, and it is where the check can be made with a plain client.
3. **Verify** `curl https://<label>.<node>.internal/` with no flags, and `trust list` naming the mesh.
Then removal, which the module declares and nothing has exercised: unassigning must take the anchor
away and refresh the bundles ([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md)),
and that is the half most likely to be wrong, because it is the half nobody reaches by accident.
@@ -0,0 +1,104 @@
# 129 — resolved: the workstation trusts the mesh, and stops when told to
*Done and measured on the machine, 2026-09-30.*
## What was done
Three steps, not the one the plan expected — the module was merged and had never been registered
(see [the diagnosis](01-diagnosis.md)):
1. **Registered** `ca-trust` from the catalogue. No artifact to build: it declares a directory, a
script, a unit and a service, and no image.
2. **Assigned** it to the workstation the issue was opened on, and pushed.
3. **Verified** with a plain client, then **unassigned and pushed again** to exercise removal, then
assigned and pushed once more.
The host's own account of arriving:
```
created ca-trust.state (/var/lib/ca-trust)
created ca-trust.anchor (/var/lib/ca-trust/anchor)
created ca-trust.unit (/etc/systemd/system/mesh-ca-trust.service)
updated ca-trust.trust (mesh-ca-trust.service): boot disabled to enabled, stopped to running
```
## It works, by the check the report asked for
The report's own reproduction, with no flags and nothing installed by hand:
```
$ curl -sS -o /dev/null -w '%{http_code}' https://git.novox.internal/
200
$ openssl s_client -connect git.novox.internal:443 -servername git.novox.internal
issuer=O=Mesh Internal CA, CN=Mesh Internal CA Intermediate CA
Verify return code: 0 (ok)
$ trust list | grep -A2 Mesh
label: Mesh Internal CA Root CA
trust: anchor
category: authority
```
Four internal names, all verifying: `git` 200, `umami` 200, `keycloak` 302, `drive` 302. Before this,
every one of them was `curl: (60) … unable to get local issuer certificate (20)`.
**And the consequence the report named specifically**: git over HTTPS to the mesh's forge, which it said
had forced the working clone URL to be ssh-only.
```
$ git ls-remote https://git.novox.internal/novox/hq.git HEAD
76fbe323ea3401fcdadbf500c61bd3fa5a0a8603 HEAD
```
## Removal is symmetric, which nothing had ever shown
[ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md) says the module anchors the
authority *and takes it away again*. That half had never run. Unassigning and pushing:
```
removed ca-trust.trust (mesh-ca-trust.service)
removed ca-trust.unit (/etc/systemd/system/mesh-ca-trust.service)
removed ca-trust.anchor (/var/lib/ca-trust/anchor)
removed ca-trust.state (/var/lib/ca-trust)
```
Then: the anchor file gone, `trust list` naming no authority of the mesh's, and the plain client back to
`unable to get local issuer certificate (20)`. A machine that leaves the mesh stops trusting it, as the
record claims.
**The order is what makes it work, and is worth saying.** The service is removed *first*, so systemd
runs the unit's `ExecStop` — which is what deletes the certificate and refreshes the bundles — while the
script it calls still exists. Had the script or the state directory gone first, stopping the unit would
have had nothing to run, and the anchor would have been left behind with nothing declaring it. Nothing
in the module says this; it is the host's removal order that makes the module's symmetry real.
## What this leaves
- **Three machines of four** *(extended the same day, after the first was proven)*. `ca-trust` is now
assigned to every converged machine, and each verifies with a plain client:
```
novox mesh CA in trust store: 1 https://git.<node>.internal/ -> 200
g14 mesh CA in trust store: 1 https://git.<node>.internal/ -> 200
shanks mesh CA in trust store: 1 https://git.<node>.internal/ -> 200
```
Before, each of the two that had not been assigned it answered
`curl: (60) … unable to get local issuer certificate (20)` and held no entry for the mesh.
**`ace` is deliberately not among them.** It is the adopted machine, still carrying the
predecessor's resolver and filter, and it is not converged until the network and module-assignment
work is settled. Assigning a module to it would hold rather than run
([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)), which is
correct and is not the same as trusting anything.
- **It arrives per assignment, which is a shape worth questioning.** The report's reasoning — that
being on the private network is what makes a machine one that speaks to the mesh's names — argues
for a fact carried to every machine on the network, the way the roster and the registry trust are.
ADR 0147 chose a module assigned per machine, and this record does not reopen it; the cost is that
a machine joining the mesh trusts nothing until somebody remembers a second command.
- **`service-manager` reports `degraded` on this machine** and the module ran anyway. Worth knowing that
the capability gate passes on a degraded service manager, since a module whose whole delivery is a
unit is the kind that would be worst served by one.
- **The predecessor's authority is still in the trust store**, beside the mesh's now rather than instead
of it. The report notes that it cannot be retired while anything on the machine speaks TLS to a mesh
name; that is no longer true here, and retiring it is its own piece of work.
@@ -1,5 +1,6 @@
--- ---
status: located status: resolved
fixed-by: mesh-host 3112c88 — undeclaring gives a unit back the state it was found in, and removes only a process the mesh made
opened: 2026-09-27 opened: 2026-09-27
located-in: [mesh-host internal/apply/apply.go (remove)] located-in: [mesh-host internal/apply/apply.go (remove)]
amended-design: 02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md amended-design: 02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md
@@ -15,7 +16,7 @@ found that the host's `remove` path stops every `service` resource that is no lo
delete". `store.Orphans` matches by id alone. So any of these stops the unit: delete". `store.Orphans` matches by id alone. So any of these stops the unit:
- the module is unassigned — by mistake, or to switch it for another; - the module is unassigned — by mistake, or to switch it for another;
- the node is sent a deliberately-empty declaration ([issue 127](../127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)); - the node is sent a deliberately-empty declaration ([issue 149](../149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md));
- a later catalogue version renames the resource's `id`. - a later catalogue version renames the resource's `id`.
That is right for a service the mesh brought into being. It is wrong for a unit the mesh That is right for a service the mesh brought into being. It is wrong for a unit the mesh
@@ -64,3 +65,9 @@ something to settle in passing.
The unassign preview is partly answered — the host's plan names each unit it will stop — and the The unassign preview is partly answered — the host's plan names each unit it will stop — and the
controller's side is left open. controller's side is left open.
## Closed
*2026-09-29, in a grooming pass rather than by whoever fixed it.* Undeclaring no longer stops a unit the mesh only reloaded or only kept running. Found by
reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was
not re-verified on a machine, and this record says so rather than implying a run that did not happen.
@@ -68,3 +68,31 @@ container runtime's shape, not a choice; the answer is to recreate, which is wha
that would rather re-read a roster from a file can already ask for one as a fact that would rather re-read a roster from a file can already ask for one as a fact
([ADR 0120](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md)) and restart on ([ADR 0120](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md)) and restart on
it. it.
## The container was umami, and it had its own record (2026-09-30)
The container in the table above is umami, and its symptom had already been filed three days earlier as
[issue 118](../118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md) — the store
answering the dial and timing out the query, a restart loop, and a public name answering `502`. Neither
record noticed the other: 118 was reasoning about path-MTU and conntrack on the network between the
container and the store, which is what a stale name looks like from inside the container.
118 is resolved by this record's fix and says so. Noted here because a symptom filed twice, diagnosed
once, and closed in one place is how a repository comes to disagree with itself.
## What replaced this fix (2026-09-30)
The fix here — putting the mesh's names into the digest the host compares, so a container whose names
moved is recreated like one whose image moved — worked, and cost more than it was worth. It made the
roster part of every container's identity, so one name moving replaced every container in the mesh: a
module assigned on one machine restarted the store, the registry, the edge and mail on another
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)).
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) goes at
the cause this record only described: **a copy taken at creation is stale the moment the roster moves,
and detecting that is not as good as not copying.** A container resolves the mesh's names through its
machine's resolver, at the moment it asks, so the fault this record reports cannot occur rather than
being noticed a restart later.
Recorded here because this is where somebody arrives to find out why the digest carries names, and the
answer is that it did, for two days short of a month, and stopped.
@@ -1,9 +1,9 @@
--- ---
status: open status: resolved
opened: 2026-09-28 opened: 2026-09-28
located-in: [mesh-controller internal/catalogue] located-in: [mesh-controller internal/catalogue/declaration.go (composeName took the consumer's own name as the internal domain)]
fixed-by: fixed-by: mesh-controller PR 163 — the internal name composes under the node whose proxy serves the route (ADR 0151, 2026-09-30)
amended-design: amended-design: 03-DESIGN/01-to-be/08-connectivity.md
--- ---
# 139 — An internal route name resolves to the consumer's node, not the one that serves it # 139 — An internal route name resolves to the consumer's node, not the one that serves it
@@ -49,3 +49,15 @@ resolve whether or not anything answers.
and if so, is `route` still one mesh-wide provision or a node-scoped seat with a mesh-wide fallback? and if so, is `route` still one mesh-wide provision or a node-scoped seat with a mesh-wide fallback?
- What certifies the name in either case? The certificate is obtained by whoever terminates TLS, and - What certifies the name in either case? The certificate is obtained by whoever terminates TLS, and
that is the question above in another form. that is the question above in another form.
## Answered (2026-09-30)
[ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md):
the internal name is composed under the node that serves the route — the machine the request arrives
at — because `<x>.<node>.internal` means *goes to that node* and nothing else. The public name stays
the consumer node's, which is where the operator put it. The first question is answered that way; the
second, a per-node route holder, is a decision about seats and is left where it is; the third is
unchanged, since the proxy that terminates the name is given it and certifies it.
mesh-controller PR 163 carries it. On this mesh every route is served beside its module, so no name
changed; the controller's tests hold the case where it would.
@@ -1,5 +1,5 @@
--- ---
status: located status: resolved
opened: 2026-09-28 opened: 2026-09-28
located-in: located-in:
- mesh-controller internal/catalogue/manifest.go - mesh-controller internal/catalogue/manifest.go
@@ -7,7 +7,7 @@ located-in:
- mesh-controller internal/catalogue/declaration.go - mesh-controller internal/catalogue/declaration.go
- mesh-controller examples/route-proxy - mesh-controller examples/route-proxy
- mesh-catalog (every routed module manifest) - mesh-catalog (every routed module manifest)
fixed-by: fixed-by: mesh-controller bdf965d (a module names its endpoints) and c68d3a7 (an assignment configures an endpoint as one thing) — the filter, the proxy's names and both authorities now read one statement
amended-design: 03-DESIGN/01-to-be/08-connectivity.md amended-design: 03-DESIGN/01-to-be/08-connectivity.md
--- ---
@@ -0,0 +1,47 @@
# Resolution
*2026-09-29.*
**Built, and this record did not say so.** The issue was written on 2026-09-28 and answered the same
week by two commits in `mesh-controller`; nothing came back to close it, so the mesh's own account of
itself said for a day that reach was declared nowhere while the code read it in three places.
- `bdf965d` — *a module names its endpoints, and a route names the one it serves*. `listens[].name`
is the endpoint; a route contribution names the endpoint rather than repeating a port.
- `c68d3a7` — *an assignment configures an endpoint as one thing*. The `endpoints` settings key, per
node, by endpoint name: `{"endpoints": {"ssh": {"port": 20134, "reach": "public"}}}` — port, label
and reach in one block, which is what [ADR 0138](../../02-DECISIONS/0138-an-assignment-binds-an-endpoint-and-says-how-far-it-reaches.md)
asked for and what [ADR 0046](../../02-DECISIONS/0046-a-module-configuration-is-its-assignments-not-its-manifest.md)
said configuration is.
## The three readers, which is what the issue was about
The complaint was that the per-node source override had exactly one caller. It now has three, and
they are the three mechanisms reach was decided to settle at once:
| reader | what it does with it |
|---|---|
| the filter | `Reaches` turns each endpoint's reach into the rule for its machine port |
| the proxy's names | `composeName` composes the public name, the internal name, or both — and a name nobody asked for is not composed |
| the authorities | the proxy certifies only names it was actually given, each from its own authority, through two host policies rather than one |
**A routed endpoint keeps the manifest's port**, which is ADR 0138's own insight and older than it
([ADR 0045](../../02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)): the
proxy is how it is reached, so `public` there asks for a public *name*, not an open port.
**An endpoint that is not routed is reached and never named.** Git over ssh is that case — the one
the issue said the model could not express — and it is now the ordinary one.
## How it is checked
`internal/catalogue/endpoints_setting_test.go`: a block says port, label and reach; a block may say
only a reach; a name the module does not declare is refused; a reach outside the four values is
refused; and saying the same thing twice — once in the block, once through the older per-port keys —
is refused rather than resolved by whichever is read last. The proxy's half is `policy_test.go` and
`authority_test.go`: a name the mesh did not send is not certified, by either authority.
## What is left, and it is not this
The older keys (`ports`, `expose`, and reach keyed by port) still work beside the block. They are
what the block replaces, and retiring them is its own small change — not a gap in what reach can
say.
@@ -74,3 +74,18 @@ every future change of this shape, and it was paid today.
argues for the former. argues for the former.
- Does the same gap apply to the launcher and the units beside the binary, which are also files no - Does the same gap apply to the launcher and the units beside the binary, which are also files no
declaration names? declaration names?
## What now waits on this (2026-09-30)
[Issue 107](../107-a-declaration-carries-no-order/00-report.md) — a declaration carries no order, so a
host cannot tell an older one from a newer. Its fix adds a field to the declaration, and a host refuses a
declaration carrying a field it does not know, **whole**. So it is a flag day: every host upgraded, then
the controller.
With nothing delivering a host, that means placing a binary by hand on every machine and unparking the
adopted one, and a machine missed in the sequence is a machine the mesh cannot send anything to at all.
107 is held for that reason rather than for anything in its own diagnosis.
**This is what makes a declaration field cost a rollout instead of an expedition**, which is a use for
this record beyond keeping machines current: it is the thing standing between the mesh and its own
protocol evolving.
@@ -0,0 +1,185 @@
# 142 — the mesh can now build and publish its own host
*2026-09-30. Both of ADR 0141's named gaps are closed; the host is not yet placed by the mesh.*
## What ADR 0141's insight named, and what each cost
> The remaining work is a way to build the host and a way to name a version in a path, and until both
> exist nothing delivers a version and every machine takes the fallback.
**1. Nothing could compile it.** The toolchain list was a closed set of typescript and python, and its
own warning — every language is another implementation of the contracts modules share, so adding one
commits to keeping N implementations in step — does not attach to Go. Go is how the host, the control
plane and the builder are written, and none of them is a module in that sense: the host is what
*applies* modules.
Closed by a Go toolchain naming a new base module, `mesh-tools-go`. Named and not pinned, so the mesh
answers with the copy it holds and moving compiler is a build rather than an edit to the control
plane's source ([ADR 0044](../../02-DECISIONS/0044-a-public-name-is-provisioned-like-any-capability.md),
[0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md)).
**2. A version could not reach the path.** An archive named a fixed path and nothing interpolated the
build into it, so nothing could ask for `…/versions/<version>/`.
Closed by `${version}` in any value of a resource that uses an archive or a bundle. **The version is
the artifact's digest, short, and not the commit**: two builds of one commit are meant to be the same
bytes — the toolchains are `-trimpath` for that — so a content-addressed version means an unchanged
build resolves to the path it already had, where a commit-named path would move for an identical binary
and recreate everything reading it.
## Three more things were in the way, and none was in the record
Found by doing it, in the order they appeared:
- **`sourcesFor` turned every entrypoint into a `.ts` file.** One language's file extension, written
into the code that serves every language. The extension is the toolchain's now.
- **The output directory was left to the compiler.** `tsc --outDir` makes one; `go build -o` writes into
a directory and does not create it, failing with a message about a path rather than about a build.
Made for every toolchain, because which compilers are forgiving is not something a reader should have
to know.
- **A bundle was refused if it named what it is built from.** The reason — a bundle is the module's own
directory compiled whole — holds for an interpreted language and cannot hold for a compiled one: a Go
repository carries several commands, the host and its bootstrap among them, and "the module's own
directory" is then not a package at all. A compiled bundle may now say which package; the refusal
stands for every interpreted one.
And a fourth, in the base module itself: its first Dockerfile ran `apt-get`, and the golang image the
mesh holds is Alpine. That is
[issue 136](../136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md) in an image
rather than on a machine, and the build refused rather than a module failing later — which is the
behaviour that issue wants.
## Measured: the mesh compiles its own host and publishes it to its own registry
```
$ mesh-controller build <the host's repository>
host-arch bundle artifact-store://mesh-host/host-arch/blobs/sha256:ad62528c…
mesh-host 1, built on novox from b5196e97
```
Fetched from the mesh's registry and opened:
```
mesh-host: ELF 64-bit LSB executable, x86-64, statically linked, stripped
$ ./mesh-host help
mesh-host — the node host
```
Statically linked matters: what a machine holds is a file rather than a container, so a binary needing
a libc it did not bring is a delivery that works until a machine differs.
**The host is a module now**, with one bundle for `arch` built from `cmd/mesh-host`. A system is
required for a compiled artifact because a binary is pinned at link time so a host refuses to touch a
machine it was not built for ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)); `arch` is what all
four of this mesh's machines report themselves to be, and another system is another artifact and
another build.
## What is left
**The host module declares no resources**, so nothing places the built bundle on a machine yet. That is
the next piece and it is the one with the interesting question in it: the resource is an archive
unpacked to `…/versions/${version}/`, applied by the host that is running, and the bootstrap is not
circular because the two are different versions ([ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md)).
The host half of that mechanism — versions side by side, the newest runs, the running one stands aside
between reconciles, rollback picks a directory — is built and tested and has never had a version to
work on.
**One thing worth noticing while writing it.** The systems list exists because "the difference between
two of them is a C library, not a kernel" — and a static Go binary has no C library. So one build would
in fact run on all three. The pin is then a policy (a host refuses a machine it was not built for)
rather than a necessity, which is a reasonable thing to keep and is worth knowing is a choice.
**And a sharper one, asked as a question and worth its own record.** The artifact's system is validated
and then read by nothing: it does not reach the compiler, no machine is matched against it, and nothing
chooses between two artifacts by it. The host built here is x86-64 because the build machine is, not
because anything in the declaration said so — correct for this mesh by coincidence. That is
[issue 159](../159-an-artifacts-system-is-checked-and-then-ignored/00-report.md).
## Delivered, started, and one fact short (2026-09-30, later)
The loop closed. The launcher is delivered as a **file** resource rather than inside the archive, and
that difference is the safety: a file is written atomically — temp file, then rename — so the running
launcher keeps the inode it was started from, where an archive writes in place with truncate and would
cut the script a running shell is reading. The manifest carries a second copy of the launcher and a
test refuses any difference from `packaging/nox-mesh-host-launch`.
On the workstation, in order: the version landed, the launcher was replaced, the running host saw a
delivered version and stood aside, and after one restart of the unit the launcher started
`/usr/lib/nox-mesh-host/versions/637f65559d16/nox-mesh-host`. **A host the mesh compiled, published,
delivered and started.**
It would then have refused the first declaration it was asked to apply. The host's Makefile links in
two facts the mesh's toolchain does not, and one of them — the system it was built for — is read before
anything is applied. That is
[issue 161](../161-a-delivered-host-carries-none-of-its-link-time-facts/00-report.md), and the machine
is back on its hand-placed binary until it is answered.
**The fallback is what made that safe**, and it was not luck: the launcher runs the pinned version, or
the newest delivered one, or the one placed by hand — so moving the delivered versions aside restored
the machine in one step.
## Self-update works (2026-09-30, end of the day)
```
running: /usr/lib/nox-mesh-host/versions/093231796eb0/nox-mesh-host
mesh-controller node show shanks
host 093231796eb0
```
One machine runs a host the mesh compiled, published, delivered and started, applying declarations and
reporting the version it was delivered as. The two facts a delivered binary was missing are
[issue 161](../161-a-delivered-host-carries-none-of-its-link-time-facts/01-resolution.md) and resolved:
the system comes from the artifact, the version from where the binary sits.
**Three machines still run a hand-placed host**, and rolling each forward is one assignment and one
push. The control node is worth last.
One thing this found on the way out: an archive cannot be undeclared, and the attempt stops the machine
applying anything at all — [issue 162](../162-an-archive-cannot-be-undeclared/00-report.md). It is how
undoing the first delivery froze the workstation, and it is not specific to the host.
## Every machine self-updates (2026-09-30, evening)
```
shanks 76f4566bef3d/nox-mesh-host active
g14 76f4566bef3d/nox-mesh-host active
novox 76f4566bef3d/nox-mesh-host active
ace 76f4566bef3d/nox-mesh-host active
mesh-controller status: (no host split)
```
The last delivery was unattended on all four: the fixed host was built, pushed, each machine stood
aside exactly once for the genuinely newer version, and the delivered launcher started it — no
restart by hand. A following push that delivered nothing new was applied and reported by every
machine and stood nobody aside, which is the check
[issue 163](../163-a-delivered-host-stood-aside-on-every-push-and-reported-nothing/00-report.md) asks
for.
**Two more faults on the way, both mine, both found by reading the machine rather than the success
line.** A delivered host compared the newest delivered version against its link-time stamp rather
than the version it was running, so it stood aside on every push and — because standing aside cancels
the report — never reported again (163). And the adopted machine kept its found launcher as the
adoption rule says, so the delivery there needed a `take` before the launcher moved.
**The crossover needs one restart of the unit per machine, once.** The launcher process that was
running on each machine was the old script, executing from its own inode; a new file beside it
changes nothing until the unit restarts. Every subsequent delivery is unattended.
**Timing, measured:** on a machine, hearing a declaration to reporting it applied is about three
seconds. A push as the operator sees it takes 17–20 seconds, and the difference is the control plane
composing the declaration before it sends. A `--wait` shorter than that reads as "did not report" for
a machine that did; the three-minute default read as slowness for a machine that never would. Neither
number is a defect being chased here, and both are worth knowing before reading a push's answer.
## What this leaves
- [Issue 162](../162-an-archive-cannot-be-undeclared/00-report.md): an archive cannot be undeclared, so
the host module — and any module with an archive — cannot be unassigned, and trying stops the machine
applying anything.
- [Issue 107](../107-a-declaration-carries-no-order/00-report.md) is unblocked: a declaration field is
now a build and a push rather than an expedition.
- Three stale version directories on the workstation from the first attempts, moved aside under
`/var/lib/mesh-host/versions-held-back/`, and a backup of the adopted machine's hand-placed binary
beside its state. Both are safe to delete and are not the mesh's to delete.
@@ -4,7 +4,7 @@ opened: 2026-09-29
located-in: located-in:
- mesh-controller internal/catalogue/filtering.go (fixed for this instance) - mesh-controller internal/catalogue/filtering.go (fixed for this instance)
- mesh-controller (what status reports, and what it does not ask) - mesh-controller (what status reports, and what it does not ask)
fixed-by: fixed-by: partly — mesh-controller 1da96e8 makes the report state its own scope; nothing dials a provision yet, which is ADR 0146 and is not built
amended-design: 03-DESIGN/01-to-be/10-delivery.md amended-design: 03-DESIGN/01-to-be/10-delivery.md
--- ---
@@ -0,0 +1,64 @@
# 145 — partly resolved: the report says what it is not a claim about
*2026-09-30.*
## What was done
The sentence that was true for eleven hours now states its own scope, immediately below itself:
```
4 machine(s), all doing what they were told, all heard from, running what the mesh would send
them, and every module current with its source
That is the mesh and the machines agreeing. Nothing here dials a provision:
no grant the mesh composed has been tested, so a module unable to reach what it
requires would not appear above (04-ISSUES/145)
```
That is the whole of what this change does, and it is deliberately small. It does not check anything.
It closes the distance between *the machines are as the mesh described them* and *it works* by naming
it, and that distance is where the eleven hours went: the report was read as the second and only ever
meant the first.
Two other reports in this group now carry real information they did not
([issue 125](../125-a-hold-is-not-a-line-in-the-apply-report/00-report.md): a hold is a line in the
apply report and breaks the all-well sentence;
[issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md): the mesh knows which
host runs a machine). Neither would have caught this fault, and both were the same shape of blindness.
`printStatus` is now separated from the asking, so these words can be read by a test with no store, bus
or machine — they have been acted on and been misleading twice, which makes them worth holding still.
## What is NOT done, and why this record stays open
**Nothing dials a provision.** [ADR 0146](../../02-DECISIONS/0146-connectivity-is-checked-by-name-per-hosting-form.md)
decides how it should be done — a module on every machine serving an endpoint of its own and dialling
every other machine's, one name per hosting form, over TLS with the certificate verified. **Nothing is
built**, the module that existed was deleted, and the work was deferred deliberately by the operator.
It is not this record's to start.
So the measured fault of this issue — that the mesh can only report on itself — is unchanged. What
changed is that the report no longer implies otherwise.
## The open questions, where they stand
- *Should a grant be checked, and from where?* Answered by ADR 0146 and not built: from the position
the callers are in, by a module on every machine, per hosting form.
- *What would it cost to be wrong in the other direction?* Unanswered and important. A check that
reports a provision broken while it works trains a reader to ignore the report, which is the failure
this whole issue is about arriving by the other door.
- *What should `status` say about a machine whose modules cannot reach each other?* Answered in part:
until something checks, it says that it has not checked. What it says when a check exists is ADR
0146's to settle.
- *Is there a cheaper signal than a probe?* Still open. The affected module logged the failure 6,154
times; the mesh reads no module's logs and arguably should not, but something a module could *say*
about its own provisions would have surfaced this in minutes.
- *Does the same blindness apply to a provider that lost a consumer's grant?* Still open, and still
nothing checks it.
## One thing worth carrying forward
**The certificate half of ADR 0146 is now possible where it was not.** Its check requires an internal
name fetched over TLS with the certificate verified, and until 2026-09-30 no machine trusted the mesh's
authority at all ([issue 129](../129-nothing-makes-a-machine-trust-the-meshs-authority/02-resolution.md)).
Three of four do now. Whoever builds 0146 no longer has to solve that first.
@@ -71,3 +71,15 @@ against a suite nobody can execute.
- The lab rewrites a bundle's registry-prefixed third-party references to upstream ones for a - The lab rewrites a bundle's registry-prefixed third-party references to upstream ones for a
machine with an uplink (`test/integration/harness.ts`); the new bundle's bus reference needed machine with an uplink (`test/integration/harness.ts`); the new bundle's bus reference needed
that rule added, which is done and is not this issue. that rule added, which is done and is not this issue.
## What one of its fixes then did to a running mesh (2026-09-30)
The change that stopped the doubling — putting the stream into a push consumer's delivery subject —
is correct on a foundation being raised and fatal on a mesh that is already running: the server will
not move that subject while a subscriber is bound, and a node is bound to its declaration consumer
the whole time it is up. The control plane crash-looped on the first build that carried it.
Recorded and fixed as [issue 156](../156-moving-a-consumers-delivery-subject-stops-the-control-plane/00-report.md).
Noted here because this record is where somebody will arrive when reading why the subject carries the
stream at all, and the answer is incomplete without it: **the raise path was the only one exercised,
and it is the one path on which nothing is bound.**
@@ -82,12 +82,100 @@ managing, or genesis carries a user list that includes the first node's enrolmen
takes over from there. Both are decisions, not patches, and both belong to the genesis step that was takes over from there. Both are decisions, not patches, and both belong to the genesis step that was
deliberately left until last. deliberately left until last.
## 5 — the composed user list has to be placed by hand at genesis *(fixed)*
The account a token is the password of is **not recorded at all**: the composer names an enrolment
user for every machine with a live token, nothing minted a credential for it, and the composition
left it out as a user with no password. The comment above the issuing code already claimed
otherwise — *"the account is created before the token is handed over"* — which is how it went
unnoticed. Issuing a token now records that account, with the token's own secret as its password,
because that is the string the machine will present.
Placing it is the other half. The list reaches the machine running the bus in that machine's
declaration, which a machine that has not enrolled does not get, so at genesis it cannot arrive
that way. **The control plane composes and says what it composed** — `broker accounts`, to standard
output — and whoever is raising the machine writes it beside the bus's configuration and makes the
server re-read it. Twice, because two accounts come into existence at different moments: the
enrolment when the token is issued, and the machine's own when it enrols. A control plane that
wrote the file itself would have to know where the bus keeps its configuration and how to make it
reload, which is the module's knowledge and is what the module takes over on the first push.
With that, **a first node enrols against the bus it just raised** — measured, from bare, in the
lab.
## 6 — and is enrolled twice, keeping a credential the mesh has replaced *(fixed)*
```
mesh-controller: enrolled anchor
mesh-controller: enrolled anchor (the same second)
```
One `enrol` on the machine, two enrolments in the control plane. Each mints the node a fresh bus
password and returns it; the machine keeps the answer to the first, and the mesh keeps the hash of
the second. The machine then reconnects for ever as a user whose password the mesh rotated out from
under it — *authentication error - User "node.anchor"* on the bus, `Authorization Violation` in the
host's log, and a node that never reports.
What is ruled out: the host asking twice — it asks again only when the mesh says *try again*, and
a refused attempt is not logged as an enrolment. Redelivery by the consumer — there is one
consumer, its acknowledgement window is thirty seconds, and the handler is quick.
What is left: the client re-publishing when an acknowledgement is slow, which is what its defaults
do. That was addressed by giving the publish a message id derived from its own bytes, so the stream
discards the copy — **and the duplicate survived it**, so either the id is not reaching the stream
or the second copy is not a copy. This is where the trail stops.
Worth saying plainly: **the mint is the fragile part, not the delivery.** An enrolment answered
twice is survivable if the answer is the same both times, and it cannot be — the mesh keeps only
the hash, so a second answer is necessarily a different credential.
**Found, and it is not about enrolment at all** *(fixed)*. The bus's own counters settled it: one
message published, one held in the stream, one delivery, nothing redelivered — and the controller
enrolled the machine twice. So the handler ran twice on one delivery.
A push consumer delivers onto an ordinary subject, and **everything subscribed to that subject gets
a copy**. The controller holds a consumer called `controller` on CONTROL and another called
`controller` on EVENTS, and the delivery subject was derived from the consumer's name alone — so
both were `_DELIVER.controller`, the one process held both subscriptions, and every message from
either stream was acted on twice.
Enrolment is where it drew blood, because enrolling twice mints twice and the second credential
replaces the first. But it applied to **every report and every event the controller follows**, and
it is the kind of fault that leaves no trace: nothing is redelivered, no counter is wrong, the work
simply happens twice. The comment in the receiving code about a merge that ran the whole catalogue
five times over on 2026-09-28 is the same shape seen from the other end.
The stream is in the delivery subject now, because the pair is what identifies a consumer — the
server scopes a durable's name to its stream, and this subject was the one place that scoping was
dropped. A subscriber's permission gains the same shape, keeping the bare name so an existing
consumer keeps working until the controller's next assertion moves it.
*How it is checked:* the consumers the mesh asks for are asserted to deliver onto distinct subjects,
in the controller's own suite. Against a server it would be invisible, which is the point.
## Where it belongs ## Where it belongs
`mesh-host` (the bundle and the enrolment path) and `mesh-controller` (the certificate command, and `mesh-host` (the bundle and the enrolment path) and `mesh-controller` (the certificate command, and
the composition that cannot reach the bus at genesis). Three of the four are fixed on branches; the the composition that cannot reach the bus at genesis). Three of the four are fixed on branches; the
fourth is the genesis work. fourth is the genesis work.
## What made it slow, and what was changed so it is not
Six faults behind one another, each found by raising a machine and reading what it said. What cost
the most was not the faults:
- **Every bed's own instructions named the bundle that cannot work**, so the first three attempts
ended in a control plane crash-looping on a missing bus. They name the working one now.
- **A host binary built without its system** refuses everything it is given with *this host was
built for ""*, which reads like a broken bundle. The lab's README says so.
- **`make image` in the control plane had been broken for as long as its base was pinned**: the
Dockerfile's fallback is a Go older than the module asks for, and the pipeline never saw it
because the pipeline passes the declared base in. It reads the base from the manifest now.
- **Leaving the machine standing is what answers the question.** Every finding above came from
shelling in afterwards — the host's log, the bus's log, the file the bus was actually handed —
and none from the test's own output, which says only that nothing converged. The bed takes
`MESH_LAB_KEEP`, and the README says to reach for it first.
## What it cost, for the next person ## What it cost, for the next person
Every lab bed still names `foundation-first-node.lock` in its own instructions, and that bundle Every lab bed still names `foundation-first-node.lock` in its own instructions, and that bundle
@@ -0,0 +1,75 @@
---
status: located
opened: 2026-09-29
located-in: [nothing in the mesh — the surface in use is the predecessor's, installed on the workstation]
---
# 147 — the operator's tools still dial the bus that was removed
## What was observed
Every tool call an operator makes against the mesh fails, on every machine, with the same answer:
```
AMQP not connected — cannot reach hal/mesh@novox
AMQP not connected — cannot reach hal/mesh@shanks
```
The mesh moved to one bus and the previous transport was deleted
([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md), cut over
2026-09-28). The surface an operator drives the mesh through — the tool bridge their assistant
speaks to — still opens an AMQP connection, so it cannot reach anything. Not one node answers,
including the machine the operator is sitting at.
**What that leaves.** The mesh itself is healthy: nodes are current, modules run, the bus carries
the mesh's own traffic. What is gone is the way a person asks it anything. Every question — what a
node runs, what is assigned, what a module's settings are — has to be asked by opening a shell on
the machine and running the control plane's binary inside its container, which is precisely the
path the tool surface exists to remove, and which nothing checks, records or permits.
It also silently changes how work gets done: an assistant told to use the mesh's tools finds them
dead, falls back to `ssh` and `docker exec`, and the operator discovers the fallback rather than
the fault.
## Why this is here and not a note in the knowledge base
The rule is that everything happens over the mesh's bus. A surface that cannot reach that bus is
that rule enforced by nothing — and the mesh reports itself healthy throughout, because what broke
is not a module, a node or a provision but the thing standing outside asking them questions.
The same cut-over removed the assistant's long-term memory (`recall`, `memorize`) for the same
reason, which is why lessons from the last two days were written into this repository by hand.
## What would have prevented it
- **The tool surface as a consumer of the bus like any other**, so moving the bus moves it — rather
than a separate bridge with its own connection settings that nothing resolves.
- **A check that a tool call reaches a node**, run where the mesh's other checks run. Every check
the mesh makes today is about the relationship between the mesh and a machine; none asks whether
a person can ask it anything ([issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
the same shape one level out).
## Diagnosed at once, because the answer was in the configuration
**The surface is not the mesh's.** The tool server the operator's assistant speaks to is the
predecessor's brain, installed on the workstation and started as a local process, with the
predecessor's broker URL — `amqp://…@<the control node>` — written into the assistant's own
configuration. Nothing about it is a module: it has no manifest, no assignment, no seat, no account
on the mesh's bus, and the mesh has never known it exists.
So nothing regressed. The mesh removed a transport that this program still dials, and the program
was never part of the mesh to be moved. **The mesh has a tool model** — `tools` in a manifest, the
calls a seat accepts, the subjects a module answers on, and a control-plane command that asks one —
and **nothing publishes an operator-facing surface onto it.** What an operator uses is the thing
that came before, kept alive by a URL in a file.
That is the issue, and it is larger than a broken connection: the way a person drives this mesh is
outside the mesh.
## Evidence to carry into diagnosis
- The failure text names `hal/mesh@<node>`, so the bridge is resolving a node and then dialling the
old transport.
- It fails identically for the local machine, which rules out reachability and points at the
transport alone.
- The mesh's own traffic over the new bus is unaffected: nodes report, declarations apply.
@@ -0,0 +1,37 @@
---
status: open
opened: 2026-09-29
located-in: [mesh-controller cmd/mesh-controller]
---
# 148 — a manifest outside this catalogue has no check
## What was observed
A module's manifest is validated by a **test** — `internal/catalogue`'s suite parses every manifest
in the catalogue checkout beside it and fails on one it cannot resolve. That works, and it is how
several real faults were caught before a machine saw them.
It is available to exactly one repository: this one. Somebody describing their own application in
their own repository — the case
[ADR 0037](../../02-DECISIONS/0037-where-a-module-lives.md) calls *the case that matters most* —
has no check at all. They write a manifest, register it with a running mesh, and find out whether
it is valid when the mesh refuses it, or later, when a machine applies something that resolved and
should not have.
The same record asks for the answer: **a `module check` command on the control plane's binary**, so
a manifest is checked by the tool rather than by a test that imports the tool's internals.
## What would have prevented it
Nothing prevents this; it was noticed and left. ADR 0037 named it on 2026-09-01 and the record sat
`proposed` until 2026-09-29, so the missing half was never anybody's task.
## Evidence to carry into diagnosis
- `mesh-controller/internal/catalogue` — `ParseManifest` and `CatalogueProblems` are the check, and
both are internal.
- The catalogue-wide test is `TestEveryCatalogueManifestDeclaresWhatItMounts` and its siblings; they
take a path from `MESH_CATALOG`, so the mechanism is already path-driven and not repository-bound.
- `mesh-controller module add` refuses a bad manifest at registration, which is the same check far
too late: by then it is in a running mesh's records.
@@ -1,10 +1,15 @@
--- ---
status: located status: resolved
opened: 2026-09-27 opened: 2026-09-27
located-in: [mesh-controller cmd/mesh-controller/push.go] located-in: [mesh-controller cmd/mesh-controller/push.go, mesh-controller cmd/mesh-controller/sendable.go]
fixed-by: mesh-controller sendable.go and push.go — an empty declaration is sent carrying `owns_nothing`, and the host refuses an empty body that does not carry it
--- ---
# A declaration that shrinks to empty is skipped, so the node keeps what it should drop # 149 — a declaration that shrinks to empty is skipped, so the node keeps what it should drop
*Opened as 127 and renumbered on 2026-09-29: two records were given that number on the same day.*
*The other kept it, because three documents and three source files cite it by number and nothing
cited this one but a decision and a sibling issue, both corrected with this move.*
## What was observed ## What was observed
@@ -37,3 +42,17 @@ mean "own nothing", which the host already applies correctly when it receives on
On ace, one command drops it permanently (the corrected controller never re-composes it): On ace, one command drops it permanently (the corrected controller never re-composes it):
`sudo ufw delete allow 5671`. At ace's converge it would clear on its own. `sudo ufw delete allow 5671`. At ace's converge it would clear on its own.
## Closed
*2026-09-29, in a grooming pass.* Both halves are on `main` and both name this issue.
- The control plane **sends** it: a declaration that composes to no resources goes out with
`owns_nothing`, and `push` says *sent, not skipped*.
- The host **refuses an empty body that does not carry it**, so a truncated or mis-composed
declaration can never be read as "own nothing" — which is the failure the fix had to avoid while
making the empty case expressible.
Closed by reading the code rather than by watching a machine let go of a stray resource; the record
says so rather than implying a run.
@@ -0,0 +1,52 @@
---
status: open
opened: 2026-09-29
located-in:
- mesh-controller internal/catalogue/declaration.go (contributions are emitted for an assigned module, taken or not)
fixed-by:
---
# 150 — A route is contributed before its module is taken
## What was observed
ace is adopted and runs `route-adapter` beside the predecessor's traefik (ADR 0104). The first web
module migrated there was searxng. `assign ace searxng` + `push ace` — the step that is supposed to
change nothing on an adopted machine (ADR 0100, hq 125: assign holds, take replaces) — took
`searxng.zurag.be` down:
```
https://searxng.zurag.be/ → 502
```
for about five minutes, until the module was unassigned again.
## Why
Assign held everything it found on the machine — the predecessor's `searxng` container, its
directories — exactly as designed. But the module's **route contribution** is not a resource on the
machine, so nothing held it: it reached `route-adapter` at once, which wrote
`mesh-searxng.zurag.be.yml` into traefik's file provider pointing at the mesh's assigned machine port
(`http://ace.internal:20000`) — where nothing listened, because the container that would was held.
traefik's file router for the name then won over the predecessor's docker-label router for the same
name, and the name served a dead backend.
On this occasion the window was lengthened by an unrelated failure (the machine's docker address
pools were exhausted, so the module's network could not be created), but the fault does not depend on
it: **between assign and take, every routed module's public name points at a backend the mesh has
deliberately not started.** The runbook's §6 step 2 ("check the preparation — HAL's service still
serves") is false for every routed module on a node running route-adapter.
## What the operator did
Rolled back per runbook (unassign before take), then re-ran as assign → push → wait for the node to
report its holds → take → push, keeping the window to the ~10 s between the two pushes. A workaround,
not a fix: a module whose data must move between assign and take (runbook §6 step 3) cannot shrink
its window this way.
## What would be right
A contribution from a module that is assigned but **not taken** on an adopted node should be held like
the module's resources are — withheld from the provider until take — or the provider should be told
the contributor is held and leave the predecessor's route alone. Either keeps "assign changes nothing"
true for routed modules.
@@ -0,0 +1,106 @@
---
status: resolved
opened: 2026-09-29
located-in:
- mesh-host internal/apply/apply.go (containerSpecReading hashes every `host` entry)
- mesh-controller internal/catalogue/declaration.go (withMeshNames gives every container the mesh's names)
fixed-by: mesh-controller PR 161 — the roster left every container's declaration and so its digest (ADR 0148 step 3, 2026-09-30)
---
# 151 — A new name recreates every container in the mesh
## What was observed
Migrating one small module on ace (searxng) took four routine controller actions: `node public-domain
ace zurag.be`, `assign ace searxng` + push, `unassign ace searxng` + push (a rollback, see
[issue 150](../150-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take
+ push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it
too"). novox's host then **replaced every container it runs, twice**:
```
22:46 … updated postgres.server (mesh-store): replaced; a container's configuration is fixed when it is created
22:46 … updated distribution.store (mesh-registry): replaced; …
22:46 … updated route-proxy.server (route-proxy): recreated …
22:51 … updated postgres.server (mesh-store): replaced; …
22:52 … updated gitea.server (gitea): replaced; …
22:52–22:55 mesh-vault, builder, all of mailu, keycloak, minio, mongodb, invoicing, photos, umami, …
```
Replacements per minute on novox, 22:46–22:55: 12, 8, 13, 5, 6, 8, 8, 8, 8, 4. The control plane was
unreachable twice while its own store came back through crash recovery
(`FATAL: the database system is starting up`), and dependants (umami) crash-looped until it did. None
of the replaced containers belonged to the module being migrated, or to ace.
## Why (confirmed part)
Every container the mesh runs is given the mesh's names as `--add-host` entries
(`withMeshNames`), and the host puts every entry into the container's spec digest — deliberately,
since [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md): a container left alone when the roster moved kept a five-day-old
address and restarted 2286 times. A running container cannot have its hosts changed, so a changed
digest means a replace.
The consequence is that **the roster is part of every container everywhere**: anything that adds,
removes or moves one name — a node's public domain, a routed name, an unassign — replaces every
container on every machine that carries the list. On the hub that includes the control plane's store,
the registry, the edge and mail.
## Not yet established
Which name moved in each pass. gitea's hosts after the fact list the `.internal` names and novox's
internally routed `*.novox.be` names, and **not** `searxng.zurag.be` — so it is not simply "a routed
name was added". Two passes suggest the set changed at assign and changed back at unassign; the
declarations before and after would say, and nothing on the machine records the previous one.
## Why it matters now
The node-by-node migration adds names one module at a time. On ace alone that is ~25 web modules; if
each assignment moves the roster, each is a full restart of every hub service, and a rollback is
another. The migration is paused on this.
## What has since been ruled out as a cause (2026-09-30)
The roster was also moving for a reason that was not a name changing at all: a node whose plan would
not compose had its routed names silently dropped from the roster handed to every machine, so a
briefly unreachable store withdrew and restored a name on alternating passes. That is
[issue 152](../152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md), and it is
fixed — it is what made the churn on the control node repeat every six minutes for twenty-five
minutes rather than once per operator action.
**It does not close this record.** A roster that changes for a real reason — a machine joining, a
module assigned, a public domain set — still replaces every container on every machine that carries
the list. 152 removed the false reasons; the question below is still open.
## What would be right (for diagnosis)
Keep 135's guarantee (no container runs with a stale address) without making the roster part of every
container's identity — e.g. resolve mesh names through a resolver the container asks at lookup time
rather than baked entries, or scope each container's entries to the names it actually binds.
## Answered (2026-09-30): the first of those two
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) takes
the resolver, not the scoping. Scoping was the close call and was rejected for contradicting the
standing intent that anything on the mesh can call anything on it: it would make a name reachable only
where the mesh was told in advance it would be wanted, and it would leave the roster in the digest, so
the churn returns whenever a widely-bound name moves.
Resolving at lookup time makes staleness impossible rather than detected — 109 and 135 stop being bugs
that were fixed and become a shape that does not exist — and makes a name's blast radius nothing.
**This record stays open**, because the record answers it and the code does not. Nothing may stop
copying names until a container can reach the resolver from any of the runtime's networks
([issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
which it cannot on two of four machines today. Removing the copy first reintroduces 109 and 135
silently, on a live mesh, which is how both were found. The order is in the record.
## Resolved (2026-09-30)
Step 3 landed the day 110 did. mesh-controller PR 161 stops writing the roster into any container, so
a container's digest no longer carries a name that is not its own. The controller's tests hold the
record's check — a container's declaration does not move when the mesh's roster does, and does move
when the module's own declared entries do.
The first push after the change recreated every container once, because every digest lost its host
entries at the same moment. That was the last such event: from here a name added or moved on one
machine changes no container anywhere, and the record's second check — add a routed name, watch every
other machine's apply report say nothing changed — is what the next module assignment will show.
@@ -0,0 +1,124 @@
---
status: resolved
opened: 2026-09-29
located-in:
- mesh-controller cmd/mesh-controller/plan.go (routeNamesInTheMesh skips a node whose plan will not compose)
fixed-by: mesh-controller 6c5dfd0 (PR 147)
amended-design:
---
# 152 — A node whose plan will not compose silently removes its names from every machine
## What was observed
For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's
host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on
the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail,
and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine
sat near 8. Nothing was converging: each pass ended and the next began four seconds later.
The two machines carrying no containers were not churning. They were only knocked off the bus each
time the control node re-created it, reconnected, re-heard the same declaration and applied it again
as a no-op.
## Why: the roster alternates between two values, and it is part of every container
Two consecutive declarations were compared by reading the `--add-host` entries of four containers
the moment each pass created them:
```
23:02 ace.internal drive.novox.be g14.internal keycloak.novox.be novox.internal
office.novox.be portainer.novox.be shanks.internal umami.novox.be (9 names)
23:07 … the same nine, and searxng.zurag.be (10 names)
```
One routed name — belonging to a module on another machine entirely — leaves the roster and comes
back. Because the roster is part of every container's spec digest ([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)),
each flip is a different identity for every container on the machine, and a running container cannot
have its hosts changed. So every flip replaces all of them.
**What makes it flip is a swallowed error.** `routeNamesInTheMesh` composes every node's plan to
find the names it serves, and when one will not compose it moves on:
```
plan, settings, err := planFor(ctx, open, n.Name)
if err != nil {
continue
}
```
A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does
not know", but "the mesh states these names do not exist", to every machine at once.
## Why it sustains itself
The loop closes through the control plane's own database:
1. An apply replaces `mesh-store` — the store the control plane reads — by removing the container,
so postgres comes back through crash recovery.
2. While it recovers it refuses connections: `FATAL: the database system is not yet accepting
connections / Consistent recovery state has not been yet reached` (observed, 21:08:39 UTC, every
pass).
3. `planFor` for the other machine fails against that store. `routeNamesInTheMesh` swallows it and
drops its routed name.
4. The roster changed, so all 327 resources differ, so all are replaced — including `mesh-store`,
and including the bus, which is why the host also cannot report: `applied, and could not tell the
mesh: reporting: nats: connection closed`.
5. Back to 1.
Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing
outside the machine has to be wrong for it to continue.
**It is metastable, not permanent.** It ran from 22:46 to 23:11 — five full replacements of every
container on the machine — and then stopped on its own, when one pass happened to read the store
during a window it was up, composed the same roster twice running, and found nothing to do. Load fell
from 7.8 to 1.7 and the machine returned to its five-minute idle tick.
That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not
control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or
a larger store, would not have found it. An outage that clears itself after twenty-five minutes and
five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it
is the same fault, harder to catch.
## What it is not
- Not the operator's four actions on the other machine. Those explain the first passes
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)); they were
finished seventeen minutes and three full passes before these measurements.
- Not a file that keeps changing. `/etc/hosts`, the bus's account list and the vault's export were
hashed across passes and are byte-identical, and the directory reported `mode 755 to 700` every
pass while already being `700`. Those resources are **misreported as changed** and are worth their
own question, but they are not what moves a container's identity.
- Not the lost report alone. A report that cannot be delivered explains a re-apply; it does not
explain a re-apply that finds 327 differences.
## Why it matters beyond this outage
The same `continue` makes every routed name in the mesh conditional on every node's plan composing at
the moment any machine is pushed to. One unreachable or half-migrated machine is enough to withdraw
its names from everywhere — and the withdrawal is indistinguishable, on the receiving machine, from
the operator having removed them.
The codebase already states the rule this breaks, forty lines away, about the same kind of lookup:
> A lookup failure is an error, never "not found": collapsing the two composed a declaration without
> the trust whenever the inventory hiccuped, delivered by a push that reported success.
## How it was fixed, and how the fix is checked
`planFor` now marks the two failures that really are a statement about the node — its set not
composing, and a setting that reaches nothing — and the three gatherers pass over those and only
those. Every other failure is raised, naming the machine and the read.
Five tests hold it: a set that cannot compose is marked as the node's own; a store that cannot be
read is *not*; one incoherent node still does not cost the rest their names; a roster is never
returned beside an error; and the raised failure names what could not be read.
The two sibling gatherers were audited and fixed the same way — the grant composer, which would have
withheld a consumer's credential, and the private-network membership, which would have taken a
machine off the overlay. Three other `planFor` callers were audited and left alone: they refuse or
report rather than silently withdraw, which is the safe direction.
**This does not close [issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md).**
A roster that changes for a real reason still replaces every container in the mesh. This removes the
false reasons; whether the roster belongs in a container's identity at all is that record's question.
@@ -8,7 +8,7 @@ fixed-by:
amended-design: amended-design:
--- ---
# 149 — An adopted machine's data cannot be placed where it is # 153 — An adopted machine's data cannot be placed where it is
## What was observed ## What was observed
@@ -0,0 +1,42 @@
---
status: open
opened: 2026-09-29
located-in:
- hq 02-DECISIONS/0138 (reach: internal | public | both)
- mesh-controller internal/catalogue/filtering.go (Reaches)
fixed-by:
amended-design:
---
# 154 — A machine's own network is not a reach
## What was observed
Preparing ace's modules. ace sits on a home network (192.168.1.0/24) behind a router, and several of
its services are reached **from that network by devices that will never be mesh machines**:
- mosquitto `1883` — an IoT light switch (`sonoff-office-light-switch`) and home-assistant;
- unifi `8080`/`3478 udp`/`10001 udp` — the access points' inform, STUN and discovery;
- plex `32400` — LAN streaming clients (three connected at survey time);
- home-assistant `8123`, and the resolver on the LAN address.
ADR 0138 gives an endpoint's reach as `internal` (the private overlay), `public` (anywhere) or `both`.
None of them says *this machine's own network*. The predecessor could: its unifi manifest opened
inform/STUN/discovery `from: 192.168.0.0/16, 10.0.0.0/8, 172.16.0.0/12`.
## Consequence
The only reach that includes a LAN device is `public`. While ace is adopted that is harmless — its
own firewall stays and admits the LAN — and behind NAT "anywhere" happens to mean the LAN. But:
- it states the wrong thing: an operator reading `reach: public` on an IoT broker believes it is on
the internet, and a router port-forward added later for something else makes it so;
- at `converge ace`, the mesh's filter is the sum of what it listens on (ADR 0045). An endpoint left
`internal` cuts every LAN device off at the flip; one set `public` opens it to the internet on any
machine with a public address.
## What would be right (for diagnosis)
A reach — or a source — that means the networks the machine is directly attached to (its uplink's
subnets, as the machine reports them), so a LAN-only service is declared as exactly that and the
filter can admit it without admitting the internet.
@@ -0,0 +1,58 @@
---
status: resolved
opened: 2026-09-29
located-in: [hq 00-META/checks/cycle.py]
fixed-by: hq f89aef9 (PR 192)
amended-design:
---
# 155 — Two records may share a number, and every check passes
## What was observed
On 2026-09-29 two machines opened issues against this repository within the same hour. Both read
`main` correctly and both took "the next free number", and they collided twice:
| | one machine opened | the other had already used |
|---|---|---|
| first | 147, 148 | 147, 148 on an unmerged branch |
| second | 149, 150 | 149 on an unmerged branch, 150 from renumbering the first collision |
The first collision was reconciled by hand before merging. The second was **merged into `main`**, and
`records.py`, `cycle.py` and `index.py` all reported success over a tree holding
`149-a-declaration-that-shrinks-to-empty` beside `149-an-adopted-machines-data-cannot-be-placed-where-it-is`,
and two folders numbered 150.
## Why
The number is allocated as `max(main) + 1`, and `main` lags every open pull request — seven of them
that evening. Two readers of the same `main` therefore compute the same next number, and neither is
doing anything wrong. The existing reconciliation precedent (a second record numbered 127 became 149)
assumed a single writer, which stopped being true when a second machine began filing its own findings.
## Why it matters
An issue number is how every other record cites this one — `fixed-by:`, `located-in:`, a decision
record's consequence, a commit message. Two records answering to one number is a citation that
resolves to whichever folder the reader happened to open, and the failure is silent on both sides:
the citer is not wrong, and the cited record exists.
It is also exactly the class this repository says it does not permit — a rule (`00-META/process/03-issues.md`:
"take the next free number") enforced by nothing.
## How it was fixed, and how the fix is checked
`cycle.py` now refuses a tree in which two issue folders share a leading number, and names both.
Proven by adding a duplicate and watching it fail, then removing it and watching it pass.
The colliding records were renumbered 153 and 154, in the branch that landed last — renumbering a
branch whose author is still pushing only moves the race.
**The check catches the collision; it does not prevent it.** Allocating a number still needs the open
pull requests read as well as `main`. That is a habit the check now backstops rather than one it
replaces, and playbook [03](../../00-META/process/03-issues.md) now says so at the step where the
number is taken.
[ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md) enumerates what `cycle.py`
enforces and named four things; it carries a progressive insight naming the fifth. The decision
stands — this is one more thing frontmatter and file names can carry, found by its absence.
@@ -0,0 +1,98 @@
---
status: resolved
opened: 2026-09-29
located-in:
- mesh-controller internal/broker/jetstream.go (EnsureConsumer)
fixed-by: mesh-controller e7da39d (PR 148)
amended-design:
---
# 156 — Moving a consumer's delivery subject stops the control plane, and only on a mesh that is running
## What was observed
The control plane crash-looped, every restart ending the same way:
```
mesh-controller: asserting how novox hears its declaration:
bringing consumer novox on NODES to match: nats: consumer name already in use
```
It came up on the first build of the controller in eight hours. The machines kept running what they
already held — this stops the mesh being *changed*, not the services it placed — and nothing could be
pushed, no report was consumed and no enrolment answered, for as long as it lasted.
## Why
[Issue 146](../146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md) put the
stream into a push consumer's delivery subject, because one process holding two consumers of the same
name on two streams was given one subject and acted on every message twice.
**The server will not move a push consumer's delivery subject while a subscriber is bound to it.** It
refuses with `consumer name already in use` — a message about the name, for a conflict about the
subject, which is why the trail starts in the wrong place.
A node is bound to its declaration consumer the whole time it is up. That *is* a node listening for
what it should be. So every node consumer in a mesh that is running is one the assertion cannot bring
to match — and the assertion happens before the controller serves, so it never serves.
The controller's own two consumers moved without trouble, and are on the new subject in the live mesh.
It asserts them before it subscribes, so nothing was bound.
## Why nothing caught it
The change was exercised on a mesh being raised, where every consumer is created rather than updated
and nothing is bound to any of them. On that path the code is correct. The test that would have caught
it needs a mesh that is already running: an existing consumer, a subscriber still attached, and then
the assertion.
Reproduced exactly that way before the fix — same server version, same stream shape, same consumer —
and it fails with the same words as the machine did. An earlier version of the same test unsubscribed
first and passed against the code that was crash-looping on the control node.
## What it is not
- Not the change it shipped beside. The merge that triggered this build carried
[issue 152](../152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md)'s fix and three
other commits that had never been deployed; this one is 146's.
- Not a version difference. The test server and the mesh's broker are both nats-server v2.10.29.
## How it was fixed
The consumer that works is kept, and the assertion says so instead of failing.
**Not deleted and re-made.** Re-making moves the subject, and a holder may not yet be allowed to
subscribe to the new one: the wider grant travels in the bus's user list, which the control plane
composes and a machine applies minutes later. On the live mesh the nodes are granted `_DELIVER.<node>`
and not `_DELIVER.<node>.>` — re-making would have silenced every machine in the mesh, which is worse
than the collision it was fixing and far harder to undo. That was the first fix written here, and the
permission is the reason it was not shipped.
**Not fatal**, which is what 146's change intended and did not do: the bare subject still delivers, and
collides only where one holder has two consumers of one name. A node has one.
## How the fix is checked
Two tests against a real server: a consumer with a subscriber bound keeps its subject, is reported,
and still delivers to that subscriber; a consumer with nothing bound moves, so 146's fix still applies
where the collision actually was.
## What is left
The node consumers stay on the bare subject, which is correct and not tidy. Nothing is wrong while
they do not move: one consumer per name per stream cannot collide with itself.
The controller reports each one it kept, and did, on the start that fixed this — four node consumers
and the build machine's worker, which is bound the same way and was not anticipated here:
```
consumer novox on NODES still delivers to "_DELIVER.novox" and not "_DELIVER.novox.NODES":
nats: consumer name already in use. It keeps working; the subject moves on an assertion
made while nothing is bound to it
```
**The wider grant has since landed** (2026-09-30, measured on the mesh's own broker config): every
node is now allowed `_DELIVER.<node>.>` as well as the bare subject. That was the thing missing when
this was diagnosed, and it is why re-making the consumers then would have silenced every machine.
What remains is only the second half — an assertion made while each node is detached from its
consumer — and that is its own piece of work, not a side effect of a restart.
@@ -0,0 +1,70 @@
---
status: resolved
opened: 2026-09-30
located-in:
- mesh-controller internal/catalogue/roster.go (the roster's entries for a routed name)
fixed-by: mesh-controller PR 163 — a routed name is published as itself, once, with no suffixed alias (ADR 0151, 2026-09-30)
amended-design:
---
# 157 — A routed name is published with an `.internal` alias that nothing serves
## What was observed
Every machine's hosts file carries two entries for each routed name — the name, and the name with
`.internal` appended:
```
10.10.0.1 keycloak.novox.be.internal keycloak.novox.be
10.10.0.1 drive.novox.be.internal drive.novox.be
10.10.0.1 umami.novox.be.internal umami.novox.be
```
The suffixed one resolves and is served by nothing. The proxy refuses it during the handshake, and
says so exactly:
```
http: TLS handshake error from 10.10.0.3:33480:
no public route for "keycloak.novox.be.internal" in this mesh, so no certificate is asked for
```
A client sees `curl: (35) TLS connect error ... tlsv1 alert internal error` and no peer certificate —
a server-side refusal, with nothing in it to say the name was never real.
The name the proxy does serve is `<label>.<node>.internal` — `keycloak.novox.internal` — which is what
[issue 139](../139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md) describes. So
there are two internal shapes for one service, one of them composed by appending the suffix to a name
that already has a domain.
## Why it matters
**It sends a reader to the wrong diagnosis.** Reproducing
[issue 129](../129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md) on 2026-09-30, the
first three names tried came from the hosts file, all failed with a TLS alert rather than the
verification error 129 reports, and the evidence pointed at the proxy having lost its internal
certificates — a regression that had not happened. The correct name reproduces 129 exactly. Several
minutes went into a fault that did not exist, and the only thing that distinguished the two was reading
the proxy's own log.
It is also a name in every machine's hosts file, and in every container's, that cannot be reached: the
shape the mesh is otherwise careful about — writing a name that resolves to something that does not
answer is worse than not writing it, because a connection to an address that does not answer hangs
where a name that does not resolve fails at once ([ADR 0007](../../02-DECISIONS/0007-connectivity.md),
and the same reasoning in `namesInTheMesh`).
## Where to look
The roster template renders one entry per name as `{{.FQDN}} {{.Name}}`, and `FQDN` is composed by
appending the mesh suffix to the bare name. For a node that is right — `novox` becomes
`novox.internal`. For a routed name the bare name is already a fully qualified public name, so the
composition produces `keycloak.novox.be.internal`, which is not a name anything was told to serve.
The fix is a judgement about what a routed name's internal form is, and 139 is the record that asks it;
this one is the evidence that the current answer publishes a third thing that is neither.
## Resolved (2026-09-30)
The judgement 139 asked for is [ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md):
a routed name has no mesh form. The roster now publishes it as itself, once, at the serving node's
address; the `<domain>.internal` line is gone from every machine's hosts file, and a controller test
refuses it coming back.
@@ -0,0 +1,49 @@
---
status: open
opened: 2026-09-30
located-in: []
fixed-by:
amended-design:
---
# 158 — The proxy re-reads and re-logs every route it serves, every two seconds
## What was observed
The route proxy on the control node logs the whole of what it serves about every two seconds —
measured 2026-09-30, **31 times in sixty seconds**, each line naming all 52 routes:
```
23:27:19 serving 52 route(s): autoconfig.novox.be, autoconfig.novox.internal, … www.praktijkdespiegel.be
23:27:21 serving 52 route(s): autoconfig.novox.be, autoconfig.novox.internal, … www.praktijkdespiegel.be
23:27:23 serving 52 route(s): autoconfig.novox.be, autoconfig.novox.internal, … www.praktijkdespiegel.be
```
Nothing is changing. The route set is identical every time.
## Why it matters
**It buries the only line that matters.** Between two of those entries sits the one error that explained
a failing name:
```
23:27:23 http: TLS handshake error from 10.10.0.3:33480:
no public route for "keycloak.novox.be.internal" in this mesh, so no certificate is asked for
```
One line of signal to roughly 5,000 characters of repetition, and the diagnosis it belonged to
([issue 157](../157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md)) was found by
grepping past it. A log that says the same true thing every two seconds is a log nobody reads, which is
the same failure as one that says nothing — and it is the mesh's own accuracy rule pointed the other
way: a report that is loud about the unchanging is not reporting.
Whether the *re-read* is also wasteful is secondary and unmeasured — it may be a cheap file stat. The
logging is not in question.
## Where to look
Not localised. The proxy is `mesh-controller examples/route-proxy`; whether it re-reads on a timer or on
a file watch, and whether it logs unconditionally or only on change, is the first thing to read. Saying
what changed — or saying nothing — is the behaviour wanted, and the mesh already has the rule written
down for its own reports: a log that is quiet on success and loud on failure reads as broken when it is
working, and one that is loud always reads as nothing.
@@ -0,0 +1,100 @@
---
status: located
opened: 2026-09-30
located-in:
- mesh-controller internal/catalogue/build.go (System is validated and read by nothing else)
- mesh-controller internal/builder (the compile invocation names no target)
fixed-by:
amended-design:
---
# 159 — An artifact's system is checked, and then nothing uses it
## What was observed
Asked whether the host is built for more than one architecture, 2026-09-30, having just built it
through the new Go toolchain.
[ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md) says an
artifact declares what it targets, and that one artifact per target is one build each. The manifest
layer enforces the first half strictly: a bundle in a language that compiles to a binary **must** name
a system, must name one of `alpine`, `android`, `arch`, and must not name one at all if its language is
interpreted. A manifest that gets any of that wrong is refused with a reason.
**The field is then read by nothing.** Every use of it in the control plane is in the function that
validates it. It does not reach the compiler, no machine is matched against it, and nothing chooses
between two artifacts by it.
So the compile runs with no target named and produces a binary for whatever the build machine happens
to be. The host, declared `system: arch` and built on this mesh's only build machine:
```
mesh-host: ELF 64-bit LSB executable, x86-64, statically linked, stripped
```
Correct for every machine in this mesh, which are all x86-64 Arch — and correct by coincidence rather
than by anything the declaration did.
## Why it matters
**A module declaring two systems would get two identical binaries.** Both would be published, both
pinned, both delivered, and the one sent to the machine it was not built for would fail at exec with a
message about a format — which is the shape ADR 0005's link-time pin exists to prevent, arriving
because the pin was never applied.
`android` in the list is the sharp end: it is not an x86-64 platform, and an artifact declared for it
today would be an x86-64 binary wearing the label. Nothing would say so until a machine tried to run
it.
**And the field reads as implemented.** It is required, validated against a closed list, and refused
with a careful message — every signal a manifest author gets says the mesh is acting on it. A field
that is checked and ignored is worse than one that does not exist, because the check is what persuades
you it works.
## Two things this is not
- **Not the same axis as the distribution.** `alpine`, `android`, `arch` are what a machine reports
itself to be, and the comment on the list says why: "the difference between two of them is a C
library, not a kernel". The processor is a second dimension and the manifest has no word for it at
all — so even a correct implementation of the current field would not answer the question that
started this.
- **Not urgent for this mesh.** Four machines, all x86-64 Arch, one build machine. Nothing is broken
today and nothing will be until a machine differs — which is exactly how long a field like this stays
invisible.
## The machine already says what it is
Found while filing [issue 160](../160-a-machine-says-little-about-itself-and-only-when-asked/00-report.md):
every machine reports its architecture and kernel in the same profile that carries its capabilities, and
the mesh keeps them.
```
ace | amd64 | linux
g14 | amd64 | linux
novox | amd64 | linux
shanks | amd64 | linux
```
Nothing in the control plane reads either, and `node show` prints the capabilities beside them without
printing them. **So a fix does not need a new fact from the machine** — matching an artifact's declared
system against what a machine reported is possible today, and the missing piece is only the comparison
and a compiler told what to target.
## Where to look
The compile invocation is assembled in `mesh-controller internal/builder`, and for Go it would need
`GOOS`/`GOARCH` set from the artifact rather than inherited from the build machine. That needs the
manifest to carry a processor as well as a system, or the systems list to mean both — which is a
decision, not a fix, and belongs with whoever answers whether one static binary should serve several
distributions at all.
**That last question is live.** The Go toolchain builds statically, so a single binary has no C library
to differ about and would in fact run on Alpine and Arch alike. The per-system pin is then a policy — a
host refuses a machine it was not built for — rather than a technical necessity, and worth knowing is a
choice.
## How a fix is checked
An artifact declared for a system the build machine is not produces a binary for that system, shown by
reading the file rather than by the build reporting success; and two artifacts declared for two systems
do not have the same digest.
@@ -0,0 +1,92 @@
---
status: open
opened: 2026-09-30
located-in:
- mesh-host internal/profile (what a machine collects about itself)
- mesh-controller internal/inventory (what the mesh keeps of it)
- mesh-controller cmd/mesh-controller (node show, which displays a part of it)
fixed-by:
amended-design:
---
# 160 — A machine says little about itself, and only when the mesh asks it something
*Filed as a to-do rather than a fault: nothing is broken by it today. `open` is the status this
repository has for "written down, nobody has started" — there is no `todo`.*
## What is wanted
**When a machine joins, the mesh should collect as much as it reasonably can about it**, and refresh
that daily or thereabouts. It already asks what the machine *can do*; what it is made of is the same
question one level down, and the mesh has no habit of asking it.
## What is already collected, which is more than it looks
A machine reports, and the mesh keeps:
| | |
|---|---|
| eight capabilities | container-runtime, firewall, graphical-session, overlay, package-manager, privileged, seat, service-manager — each with a version or a reason it is absent |
| **its architecture and kernel** | in the same profile, beside the capabilities |
| which links face outside it | read from its own routing table on every apply |
| the version of the host running on it | added 2026-09-30 |
| what it found and is holding | on an adopted machine |
| what is reachable on it | every listening socket and published port, on an adopted machine |
| the firewall it was found with, and the tunnel it carried | on an adopted machine |
**The architecture and the kernel are already there and nothing reads them.** Measured on the mesh:
```
ace | amd64 | linux
g14 | amd64 | linux
novox | amd64 | linux
shanks | amd64 | linux
```
`node show` prints the capabilities and not these two. No code in the control plane reads either.
## What is missing
**The facts.** Nothing is collected about memory, disk, the processor beyond its architecture, the
distribution and its version, whether the machine is virtual or physical, its uptime, its timezone, how
many cores it has. Several of those are what somebody actually wants when deciding where a module
should go, and the comment on module placement already imagines them: *"`seat: card1-DP-1`, an
architecture, an amount of memory"*.
**The refresh.** A machine publishes what it knows about itself **after an apply**, and the five-minute
reconcile publishes nothing. So the mesh's picture of a machine is as old as the last time it sent that
machine something — found the same day while checking host versions, where three machines read as
having said nothing until each was pushed
([issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/01-resolution.md)). A daily refresh is
the same missing mechanism: a machine saying something on its own schedule rather than only when spoken
to.
**Somewhere to read it.** Whatever is collected has to be visible, or it joins the architecture in
being true and unread.
## The pattern this is the third instance of
Three times in one day the mesh was found to be collecting something and reading it nowhere:
- what an adopted machine holds — reported since the adoption work, and no surface counted it until
[issue 125](../125-a-hold-is-not-a-line-in-the-apply-report/01-resolution.md);
- the host version — reported for a week, and the control plane's own copy of the report did not have
the field ([issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md));
- the architecture and kernel — collected, stored, read by nothing, found today by being asked whether
the host is built for more than one processor
([issue 159](../159-an-artifacts-system-is-checked-and-then-ignored/00-report.md)).
**Collecting is the easy half and it is the half that gets done.** Whatever this issue adds should say,
in the same breath, which surface shows it — or it will be the fourth.
## Why it is not urgent
Every machine in this mesh is `amd64` and every one reports so. One architecture is enough for now, and
that is a decision rather than an oversight: a second processor is a second build of every component,
and nothing needs one.
## How a fix is checked
A machine's own account of itself is visible in one place; it names at least what is listed as missing
above; the newest of it is no older than a day on a machine nobody has pushed to; and a machine that
cannot determine one of them says so rather than reporting a zero.
@@ -0,0 +1,95 @@
---
status: resolved
opened: 2026-09-30
located-in:
- mesh-host cmd/mesh-host/main.go (version and builtFor, both set at link time)
- mesh-controller internal/builder (the toolchain, which deliberately takes nothing from the module)
fixed-by: mesh-controller (the system stamp, and one linker flag rather than two), mesh-host (the version read from the path) — verified on a machine 2026-09-30, 01-resolution.md
amended-design:
---
# 161 — A host the mesh built carries none of the facts its Makefile stamps in
## What was observed
*2026-09-30, on the workstation, having just made the host self-updating.*
The mesh compiled the host, published it, delivered it and the launcher started it. It ran, read the
machine correctly, and **would have refused the first declaration it was asked to apply.**
The host's own Makefile links in two facts:
```
LDFLAGS := -s -w -X main.builtFor=$(SYSTEM) -X main.version=$(VERSION)
```
The mesh's Go toolchain links in neither, on purpose: a toolchain accepts nothing from the module,
because anything a module could override there it would be writing a Dockerfile to override
([ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md)). So a
delivered host has `builtFor = ""` and `version = "development build"`.
**`builtFor` empty is the one that bites.** Before applying anything, the host asks which system it
was built for:
```go
sys, err := system.For(builtFor)
```
and that answers, for an empty name:
```
this host was built for "", which is not a system it knows. Built hosts are: …
```
It is called before any resource is applied, so the failure is in the safe direction — the machine is
not half-configured. It is still a host that cannot do its job, and nothing about it looks wrong: the
unit is active, the link to the bus is up, and the log says it is hearing what the node should be.
Measured: after the crossover the machine logged nothing further, where the previous host had written
a reconcile line every five minutes.
## Why this was found rather than reported
Nothing reports it. The host does not check its own stamps at start, the mesh does not ask, and the
declaration that would fail is the same declaration that would deliver a fix — so **a machine in this
state cannot be repaired by the mesh.** It was restored by moving the delivered versions aside and
letting the launcher fall back to the hand-placed binary, which is the fallback working exactly as
designed.
## What the records already say about half of it
[ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md) settles the
version and its answer is not implemented:
> A component's version comes from where it sits, not from its linker. It is unpacked into a directory
> named for its version, so it can read its own version from its path. The stamp goes, and with it the
> need for a build to know what it will be called.
That is exactly right and would also fix what the mesh reports: a delivered host would say
`637f65559d16` rather than `development build`, and
[issue 087](../087-the-controller-cannot-tell-a-host-is-too-old/00-report.md)'s host comparison would
mean something for delivered hosts.
**The system pin has no answer yet**, and it needs one before any mesh-built host can apply anything.
The tension is real: the target is a property of the artifact and 0142 says so, but a toolchain that
passed it would be linking a value into a variable whose name belongs to the module — which is the
coupling the toolchain exists to avoid. Candidates, none decided:
- the path carries it as well as the version, so the host reads both from where it sits, as 0142 does
for the version;
- the bundle carries a small file beside the binary saying what it was built for, written by the
builder from the artifact's declaration;
- the host stops being pinned at link time and refuses on a fact it reads from the machine instead —
which changes what [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) decided and is the biggest of
the three.
## What is true in the meantime
Self-update works end to end and is one fact short of usable: the mesh builds the host, publishes it,
delivers it to a machine, the running host stands aside, and the launcher starts the delivered one. The
machine is left on its hand-placed binary until this is answered, which is one command to undo.
## How a fix is checked
A host the mesh built and delivered applies a declaration on a machine, shown by the machine's own
reconcile line; and it reports a version that names the build it came from rather than a placeholder.
@@ -0,0 +1,59 @@
# 161 — resolved: a host the mesh built runs a machine
*2026-09-30. Measured on the workstation.*
```
running: /usr/lib/nox-mesh-host/versions/093231796eb0/nox-mesh-host
agent: active
reconciles in the last six minutes: 1
mesh-controller node show shanks
host 093231796eb0
```
A binary the mesh compiled, published to its own registry, delivered over the bus, started by the
launcher, applying declarations, and reporting a version that names the build it came from.
## The two facts, and where each now comes from
**The system it was built for comes from the artifact.** ADR 0142 already made the target a property
of the artifact rather than of the recipe, so the toolchain names the variable it fills and the
artifact supplies the value. It is the one thing a toolchain takes from a module, and it is stated
rather than inferred.
**The version comes from where the binary sits**, which is what
[ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md) decided and
nothing had implemented: a delivered host reads the directory it was unpacked into. A host placed by
hand keeps its link-time stamp, which is the honest answer for one the mesh did not deliver — and is
every other machine today.
## Two mistakes on the way, both found by reading the output
**A repeated flag is not a merged one.** The stamp was appended as a second `-ldflags`, and the Go
command takes the last and drops the first. The binary gained its system and lost `-s -w`: 12.2MB
against 8.5MB, with its debug info. The comment I had written said the linker "accepts and merges"
them. It does not. Linker flags are the toolchain's own list now, composed into one flag, and a test
refuses a compile line that carries `-ldflags` itself.
**The delivered binary was named after its package.** `cmd/mesh-host` builds `mesh-host`; every
machine runs `nox-mesh-host`, which is what the launcher looks for inside a version. The first
delivery landed, reported `created … 1 file(s)`, and was invisible. An artifact says what its
executable is called now.
Both were caught by listing the directory and reading the binary rather than believing the line that
said it worked.
## What this cost while it was wrong, and what saved it
A delivered host that cannot apply is a machine the mesh cannot repair, because the declaration that
would fix it is the declaration it cannot apply. The workstation was restored by moving the delivered
versions aside so the launcher fell back to the hand-placed binary — **the fallback in the launcher,
working exactly as designed**, and the reason this was an inconvenience rather than an expedition.
It also loops if you are not careful: the working binary applies, delivers a version, stands aside,
and the broken one starts. Stopping the unit while the fix was built was the way through.
## What is left
**Three machines still run a hand-placed host.** Rolling them forward is one assignment and one push
each, and the control node is worth doing last and watching.
@@ -0,0 +1,56 @@
---
status: located
opened: 2026-09-30
located-in: [mesh-host internal/apply (no removal for an archive)]
fixed-by:
amended-design:
---
# 162 — An archive cannot be undeclared, and trying stops the machine applying anything
## What was observed
*2026-09-30, unassigning the host module from the workstation to undo a delivery.*
```
holding this machine: 0 applied, and map[apply:applying "mesh-host.next": no way to remove a "archive"
0 resource(s) were applied and remain; everything was attempted, so what is not listed as failed was done.
```
**Nothing was applied at all** — not the archive, not the other forty resources that had nothing to do
with it. The machine stopped reconciling and stayed that way until the module was assigned again.
## Why it matters
Every other resource kind can be taken away. A file is removed and what was found under it is put
back; a container is stopped and removed; a unit is given back the state it was found in
([ADR 0118](../../02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md)). An
archive has no removal at all, so:
- **a module with an archive can never be unassigned** — the attempt fails for ever;
- **the failure takes the whole apply with it**, so the machine applies nothing else either, and one
unassignable resource is a machine frozen against every other change;
- it is silent in the mesh's terms: the push reported sent, and only the machine's own journal says
what happened.
The host module is the obvious case and not the only one. An archive is for what inlining cannot
serve — a theme, an icon set, a tree of configuration — and any module using one is in the same
position.
## What the right answer probably is, and the question in it
The other kinds answer this by remembering what they found. An archive unpacks many files into a
directory the mesh did not necessarily create, so removal has a real question in it: **remove what the
archive put there, or remove the directory?** The first needs the applier to have recorded the file
list; the second would delete whatever else lives there — and for the host's own versions directory,
that is every other delivered version.
Recording what was unpacked is the answer that matches how the rest of the host behaves, and it is
what [issue 126](../126-a-volume-path-is-not-in-the-spec-comparison/00-report.md) and ADR 0118 already
argue for elsewhere: the mesh gives back what it found.
## How a fix is checked
A module with an archive is assigned, pushed, unassigned and pushed again; what the archive put on the
machine is gone, anything that was in the directory beforehand is still there, and the apply that
removed it applied everything else in the same declaration.
@@ -0,0 +1,56 @@
---
status: resolved
opened: 2026-09-30
located-in: [mesh-host cmd/mesh-host/main.go (the successor check after an apply)]
fixed-by: mesh-host PR 58 — the check asks with the running version, read from the binary's path, not the link-time stamp
amended-design:
---
# 163 — A delivered host stood aside on every push, and reported nothing
## What was observed
*2026-09-30, rolling the mesh-built host onto the last two machines.*
Every push to a machine running a delivered host produced, in order:
```
host 093231796eb0 is delivered; standing aside so the launcher runs it
applied 333 resource(s)
applied, and could not tell the mesh: reporting: context canceled
nox-mesh-host-launch: the host exited cleanly; starting it again
nox-mesh-host-launch: running /usr/lib/nox-mesh-host/versions/093231796eb0/nox-mesh-host
```
— for the version it was **already running**. It restarted itself on every push, for ever, and the mesh
never received a single report from it: `node show` kept the version from before the crossover, and
the operator's push waited its full three minutes for an answer that was never coming.
Read as healthy throughout: unit active, bus link up, "hearing what this node should be".
## Why
After an apply the host asks whether a newer host has been delivered than the one running, and the
question was asked with the **link-time version stamp**. Since
[issue 161](../161-a-delivered-host-carries-none-of-its-link-time-facts/01-resolution.md) a delivered
host's version comes from where it sits and its stamp is `development build` — so the comparison never
matched the newest delivered version, and "a newer host is waiting" was always true.
Standing aside cancels the context the report is published with, so the report was lost on every one
of those applies. Two faults from one wrong argument.
The change that moved the version to the path was applied to the report and to the known-good record,
and not here. Half a change, and the half left behind was the one that decides whether to exit.
## Why the three-minute wait made it invisible
The push's `--wait` timing out read as *slow*. It was not slow: **the report was never going to arrive.**
The operator put it exactly: *if you don't get a response in five seconds, something is wrong.* A wait
long enough to absorb a machine's whole apply is a wait long enough to hide that the machine never
answered.
## How it is checked
A machine running a delivered host is pushed a declaration that delivers nothing new; it applies,
reports, and does not stand aside. A machine running a delivered host is pushed a genuinely newer
version; it stands aside once, and the next push it does not.
@@ -0,0 +1,78 @@
---
status: resolved
opened: 2026-09-30
located-in:
- mesh-catalog modules/mailu (eight containers name Mailu's own resolver, and one of them binds a mesh name)
- mesh-catalog modules/dnsmasq (dropped the DNSSEC bit its upstreams set)
fixed-by: mesh-catalog PR 178 (mailu-admin uses the machine's resolver) and PR 179 (the machine's resolver passes the DNSSEC bit down) — 2026-09-30, the same afternoon
amended-design:
---
# 171 — A module that names its own resolver knows no mesh name
## What was observed
The afternoon [ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md)
landed — no container is given the mesh's names any more; it asks the machine's resolver — the mail
system's admin container began logging, 523 times in three minutes:
```
psycopg2.OperationalError: could not translate host name "novox.internal" to address: Name does not resolve
```
Mail was accepted on every port and the web front answered; the admin and the spam filter beside it
were unhealthy, and anything that needed the database — a mailbox change through the API, the spam
filter's domain list — failed. Found by the operator asking whether mail was back, forty minutes in.
Mailu ships its own resolver, an unbound in a container, and every other Mailu container is told to
use it — the module carries `dns: [192.168.203.254]` on eight containers. That resolver recurses from the
root and knows nothing under `.internal`. Until that afternoon the admin container had the database's
name anyway, because the mesh wrote every name into every container at creation; the copy was the only
reason a container behind its own resolver could reach anything by a mesh name, and nobody knew it was
load-bearing.
**Removing the override was not enough.** Given the machine's resolver instead, the admin refused to
start: `Your DNS resolver at 127.0.0.11 isn't doing DNSSEC validation`. Mailu checks, at start, that
its resolver returns the Authenticated Data bit for a signed name. The mesh's resolver forwards to two
upstreams that validate and set the bit, and dropped it on the way down — dnsmasq does unless told
otherwise. Mailu's own unbound has no hook to forward a zone elsewhere, so it could not be taught the
mesh's names either.
## Why it matters beyond this instance
**A container with a resolver of its own has opted out of the machine's, and nothing says so.** 0148
made the machine's resolver load-bearing for every container; a `dns` on a container is a quiet
exception to that, and the exception used to be papered over by the copy the record removed. The
manifest field reads like a preference and is a decision about whether mesh names exist inside the
container.
**A resolver that forwards to validating upstreams and hides the fact is less useful than it could
be**, and the first program to check found out.
**The mesh reported nothing.** Every container ran; the failing one accepted connections; the report
was about bytes. It is [issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
again, and the check that would have caught it is the same unbuilt one.
## What was done
- The one Mailu container that binds a mesh name — the admin, through the database it is granted —
no longer names Mailu's resolver and uses the machine's, like every container without a `dns` of
its own (mesh-catalog PR 178). The other seven keep unbound: the spam filter needs a validating
resolver for its blocklist lookups, and none of them asks for a mesh name.
- The machine's resolver passes the DNSSEC bit down from its upstreams, `proxy-dnssec` (PR 179). It
does not validate itself; the trust is the upstream's and the path to it, as a forwarding resolver's
always was, and the configuration says so.
## What checks it
The admin container's own start-up check, which is what failed, and the mesh's status once it reads
healthy. A container-level check that a mesh name resolves from inside every declared container is
the one 110 and 145 both ask for and is not built.
## Open questions
- Should a container's `dns` be refused, or made to say what it gives up? A module that names its
own resolver and binds a mesh name is a contradiction the controller can see at composition — the
grant hands it a name its resolver will not answer.
- Should the machine's resolver validate rather than proxy? It would cost a trust anchor on every
machine and make the resolver slower to start; proxying was enough for the one program that asked.
@@ -0,0 +1,55 @@
---
status: located
opened: 2026-09-30
located-in:
- the predecessor's terminal module (still generating the operator's ssh client blocks on every workstation)
- mesh-controller internal/catalogue (the ssh-client roster, tested and not yet a catalogue module)
fixed-by:
amended-design:
---
# 172 — The ssh client block for a machine matches one spelling of its name, and the other gets the wrong user
## What was observed
On a workstation, 2026-09-30, reported by the operator. `ssh home-server` logs in; `ssh
home-server.internal` is refused with `Permission denied (publickey)`. The operator expected the
opposite, if either: the full mesh name is the one the resolver serves.
The name is not the fault. Both spellings resolve to the machine's private address — the mesh's roster
region in the hosts file carries `<node>.internal <node>` on one line, and the resolver answers
anything under the node's name. What differs is the login: the generated client configuration has a
`Host home-server` block naming the account to log in as, and `home-server.internal` matches no block,
so ssh falls back to the operator's local username, which has no account on that machine. Spelled
`account@home-server.internal` it works.
The file is the predecessor's. `~/.ssh/config.d/mesh` says in its own header that it is generated by
the predecessor's terminal module, which only ever wrote the bare name. The mesh's own ssh-client
roster — every other machine's Host block, written as a marked region of the operator's `~/.ssh/config`
with the account the mesh knows for that machine (to-be 29) — already matches both spellings, and a
controller test holds `Host marge marge.internal`. It is composed and tested in the controller and is
not a module in the catalogue, so no machine receives it; every workstation still runs the
predecessor's generator.
## Why it matters beyond this instance
**A name the mesh serves and a name a person can use are not the same set**, and the difference is
silent. The resolver, the hosts file and the certificate authority all treat `<node>.internal` as the
machine's name; the one file that decides who you log in as does not know it. A person who learns the
mesh's name from `status` or from a certificate and types it is refused with an error that says
nothing about a missing Host block.
**It is the migration story for the operator's own tooling, arriving as a symptom.** The mesh has the
right file and does not ship it. Until the ssh-client roster is a module and is assigned to the
workstations, the predecessor's generator keeps writing a file the mesh has already superseded, and
every such file is one the mesh cannot correct.
## Open questions
- Should the ssh-client roster become a catalogue module now, assigned to every workstation, and take
the predecessor's `config.d/mesh` out of the operator's `Include`? Its content is settled; what is
not is the takeover of a file in a person's home that another generator still writes.
- Should the block match a third spelling — the machine's public name, where it has one — or is that
a different key and a different account?
- What checks it? A controller test holds the two spellings; nothing checks that the file a workstation
actually has is the mesh's rather than the predecessor's.