885 Commits
Author SHA1 Message Date
jschoubben 0cf1ad5dad Merge pull request 'Issue 172: the ssh client block matches one spelling of a machine's name' (#221) from issue/172-the-ssh-client-block-matches-one-spelling-of-a-machine into main 2026-09-30 13:25:34 +00:00
jschoubben 69a002fce3 Issue 172: the ssh client block matches one spelling of a machine's name
Reported by the operator: ssh by the bare name logs in, by the mesh name
is refused. The predecessor's generator writes the bare name only; the
mesh's ssh-client roster already matches both and is not yet shipped.
2026-09-30 15:25:30 +02:00
jschoubben 16a1a52cd8 Merge pull request 'Issue 171: a module that names its own resolver knows no mesh name' (#220) from issue/171-a-modules-own-resolver-knows-no-mesh-name into main 2026-09-30 13:20:30 +00:00
jschoubben af170e3a67 Issue 171: a module that names its own resolver knows no mesh name
Found and fixed the afternoon ADR 0148 landed: mailu-admin lost its
database behind Mailu's own resolver. Two catalogue PRs; an insight on
0148 that a container's dns is a decision, not a preference.
2026-09-30 15:20:26 +02:00
jschoubben 6c2d5f5913 Merge pull request 'Group 2 is resolved: containers resolve, nothing is copied, a route's name says where it arrives' (#219) from issue/110-resolved into main 2026-09-30 13:07:14 +00:00
jschoubben 04c9500b5b Group 2 is resolved: containers resolve, nothing is copied, a route's name says where it arrives
Issue 110's cause was not the filter: the runtime had never been told,
and the resolver dropped a query arriving on a bridge. ADR 0148 step 3
landed once it did (109, 151 resolved). ADR 0151 composes a route's
internal name under the serving node and drops the suffixed alias
(139, 157 resolved). Design 08 amended; a fact in 0148 corrected.
2026-09-30 14:56:43 +02:00
jschoubben 846c1f85f2 Merge pull request 'Issue 107 is resolved: a declaration carries its order' (#217) from issue/107-resolved into main 2026-09-30 12:13:57 +00:00
jschoubben 9eef0bd525 Issue 107 is resolved: a declaration carries its order
Hosts first, then the controller — a build and a push each, now that the
mesh delivers the host. The host refuses a lower sequence than it kept
and drains a batch by sequence rather than arrival; the controller
numbers each send under the node's hold, inside the signed bytes.

Measured: two pushes, sequence 2 in the kept declaration, counters in
the store agree, no machine reads as behind. That last one is the
subtlety: the mesh compares the digest of what it would send against
what it did, and a number changes the bytes, so the read-only comparison
composes with the last number sent rather than a fresh one.
2026-09-30 14:13:50 +02:00
jschoubben 6e08cdf3d6 Merge pull request 'Every machine self-updates, verified, and 107's gate has opened' (#216) from issue/142-self-update-on-every-machine into main 2026-09-30 11:51:53 +00:00
jschoubben 02f291a129 Every machine self-updates, verified, and 107's gate has opened
All four machines run a host the mesh built, published and delivered, the
last delivery unattended: each stood aside once for a genuinely newer
version and the delivered launcher started it. A following push that
delivered nothing new was applied and reported by every machine and stood
nobody aside.

The crossover needs one restart of the unit per machine, once, because
the running launcher executes from its own inode. Measured timing: three
seconds on the machine, 17-20 as the operator sees it, the difference
being the control plane composing before it sends.

107 is unblocked: a declaration field is now a build and a push.
2026-09-30 13:51:46 +02:00
jschoubben 9b14430d3f Merge pull request 'Issue 163: a delivered host stood aside on every push and reported nothing' (#214) from issue/163-a-delivered-host-stands-aside-on-every-push into main 2026-09-30 11:47:35 +00:00
jschoubben a4384f13d3 Issue 163: a delivered host stood aside on every push and reported nothing
Asked whether a newer host was delivered using the link-time stamp, which
every delivered host carries as 'development build' now that the version
comes from where the binary sits. Never matched, so it stood aside on
every push for ever; standing aside cancels the report, so the mesh never
heard from it. Read as healthy throughout.

The three-minute push wait made it invisible: a wait long enough to
absorb a whole apply is long enough to hide that the machine never
answered.
2026-09-30 13:47:28 +02:00
jschoubben a841e2c173 Merge pull request 'The host self-updates, and an archive cannot be undeclared' (#212) from issue/161-resolved-and-162-an-archive-cannot-be-removed into main 2026-09-30 11:16:44 +00:00
jschoubben 3c535ead31 The host self-updates, and an archive cannot be undeclared
161 resolved and verified on a machine: the workstation runs a host the
mesh compiled, published, delivered and started, applying declarations
and reporting the version it was delivered as.

The system it was built for comes from the artifact — the one thing a
toolchain takes from a module, which 0142 already allowed because the
target is a property of the artifact. The version comes from where the
binary sits, which 0142 decided and nothing had implemented.

Two mistakes on the way, both caught by reading the output rather than
the line that claimed success. A second -ldflags does not merge with the
first: the binary gained its system and lost -s -w, 12.2MB against 8.5MB.
And the delivered binary was named after its package, so the first
delivery was correct, reported success and was invisible to the launcher.

A delivered host that cannot apply is a machine the mesh cannot repair,
because the declaration that would fix it is the one it cannot apply. The
launcher's fallback is what made that an inconvenience instead of an
expedition.

162 is new and not about the host: an archive has no removal, so a module
using one can never be unassigned, and the attempt takes the whole apply
with it — the machine applies nothing else either. It is how undoing the
first delivery froze the workstation.
2026-09-30 13:16:37 +02:00
jschoubben 4cf941d858 Merge pull request 'Self-update works, and a delivered host is one fact short of usable' (#211) from issue/161-a-delivered-host-has-no-link-time-facts into main 2026-09-30 10:28:08 +00:00
jschoubben 5042ffd8d3 Self-update works, and a delivered host is one fact short of usable
The loop closed on the workstation: the version landed, the launcher was
replaced, the running host stood aside, and after one restart the launcher
started a binary the mesh had compiled, published and delivered.

The launcher goes as a file resource rather than inside the archive, and
that is the safety rather than a preference. A file is written atomically,
so the running launcher keeps the inode it started from; an archive writes
in place with truncate and would cut a script a shell is reading. The
manifest carries a second copy and a test refuses any drift from the one
in packaging.

Then it would have refused the first declaration it was asked to apply.
The Makefile links in two facts the mesh's toolchain does not, on purpose,
and one of them is the system the host was built for — read before
anything is applied, so the failure is safe and total. Nothing reports it:
the unit is active, the bus link is up, and the log says it is hearing
what the node should be.

Worse, the declaration that would fix it is the declaration it cannot
apply, so the mesh cannot repair such a machine. Restored by moving the
delivered versions aside and letting the launcher fall back, which is the
fallback working as designed.

0142 already settles the version — it comes from where the component sits,
not from its linker — and that is unimplemented. The system pin has no
answer, and the candidates are a decision rather than a fix: put it in the
path too, carry it in a file beside the binary, or stop pinning at link
time at all, which is 0005's to change.
2026-09-30 12:28:01 +02:00
jschoubben 4f9dc3906e Merge pull request 'A machine says little about itself, and only when the mesh asks it something' (#209) from issue/160-what-a-machine-says-about-itself into main 2026-09-30 08:36:48 +00:00
jschoubben a2542e51f8 A machine says little about itself, and only when the mesh asks it something
Filed as a to-do. Nothing is broken by it: every machine here is amd64 and
reports so, and one architecture is enough for now.

When a machine joins, the mesh should collect what it reasonably can about
it and refresh that daily. It already asks what a machine can do; what it
is made of is the same question one level down.

More is already collected than it looks — eight capabilities, the links
that face outside, the host version, and on an adopted machine what it
holds, what is reachable and the firewall and tunnel it was found with.
The architecture and the kernel are in there too, and nothing reads
either: measured, all four machines report amd64 and linux, and node show
prints the capabilities beside them without printing them.

Missing: memory, disk, the processor beyond its architecture, the
distribution and its version, virtual or physical, cores, uptime. Several
are what somebody wants when deciding where a module goes, and the
placement code's own comment already imagines them.

Also missing: the refresh. A machine publishes after an apply, and the
five-minute reconcile publishes nothing, so the mesh's picture is as old
as the last push. Same mechanism 087 wanted.

This is the third thing in one day found to be collected and read nowhere,
after held resources and the host version. Whatever gets added should say
in the same breath which surface shows it, or it will be the fourth.

159 gains the note that the architecture is already reported, so matching
an artifact to a machine needs no new fact — only the comparison and a
compiler told what to target.
2026-09-30 10:36:40 +02:00
jschoubben b19b29cd3a Merge pull request 'An artifact's system is checked and then nothing uses it' (#208) from issue/159-an-artifacts-system-is-checked-and-ignored into main 2026-09-30 08:30:52 +00:00
jschoubben 4bf4fe2d06 An artifact's system is checked and then nothing uses it
Asked whether the host is built for more than one architecture. It is
not, and the reason is worse than a missing feature.

A bundle in a compiled language must name a system, must name one of
alpine, android or arch, and is refused with a careful message if it gets
that wrong. The field is then read by nothing: it does not reach the
compiler, no machine is matched against it, and nothing chooses between
two artifacts by it. The compile runs with no target named and produces a
binary for whatever the build machine happens to be.

The host is x86-64 because the build machine is, not because the
declaration said so. Correct for this mesh by coincidence — four machines,
all x86-64 Arch.

A module declaring two systems would get two identical binaries, both
published and both pinned, and the one sent to the machine it was not
built for would fail at exec. android is the sharp end: not an x86-64
platform, and an artifact declared for it today would be an x86-64 binary
wearing the label.

A field that is checked and ignored is worse than one that does not
exist, because the check is what persuades you it works.

Also noted: the processor is a second dimension the manifest has no word
for, so even implementing the present field would not answer the question
that found this. And since the Go toolchain builds statically, one binary
would run on all three systems anyway — so the pin is a policy rather
than a necessity, which is a decision and not a fix.
2026-09-30 10:30:45 +02:00
jschoubben 5e12081774 Merge pull request 'Issue 142: the mesh compiles its own host and publishes it to its registry' (#207) from issue/142-the-mesh-can-build-its-own-host into main 2026-09-30 08:05:12 +00:00
jschoubben 9ad64881a9 Issue 142: the mesh compiles its own host and publishes it to its registry
Both of the things ADR 0141's insight named as remaining are built. A Go
toolchain based on a new mesh-tools-go module, so the compiler is named
and not pinned; and ${version} in any value of a resource that uses an
archive or a bundle.

Measured rather than asserted: the mesh built the host through its own
toolchain, published it to its own registry, and the bundle fetched back
out is a statically linked stripped binary that runs and says it is the
host.

The cost was larger again than 0141's note said. Three more things in the
path assumed one language or one shape — an entrypoint became a .ts file
whatever the language, the output directory was the compiler's to create,
and a bundle was refused if it named what it is built from — and a fourth
was in the base image, which is Alpine where the first Dockerfile ran
apt-get. That last one is issue 136 in an image, and the build refused
rather than a module failing later.

The version in a path is the digest, not the commit: two builds of one
commit are the same bytes, so a content-addressed version keeps the path
an unchanged build already had.

Still nothing delivers a version to a machine. The host module declares no
resources, so the bundle sits in the registry and no machine is asked to
take it. 0141 carries the insight and 142 the account.
2026-09-30 10:05:04 +02:00
jschoubben 266ade6e28 Merge pull request 'Issue 087: what I shipped first said the opposite of the truth' (#206) from issue/087-a-commit-has-no-order into main 2026-09-30 07:29:53 +00:00
jschoubben a794ef9a3f Issue 087: what I shipped first said the opposite of the truth
It reported "N machines run an older host than another" by comparing
versions as strings. A host reports its version as a commit, and commits
have no order. On the live mesh it named the three machines running the
NEWER host as the ones behind — ced54d4 sorts above 04a27ca and that is
all it means.

The code even carried a caveat saying versions compare as strings and that
this "is enough for the timestamps and commits this mesh uses". That was
the error, written down and not noticed: enough for timestamps, meaningless
for commits, and the mesh reports commits.

It now reports the split and claims no ordering, which is more useful as
well as more honest — the reader sees who is on which side, and that is
what decides whether a field can be sent. Ordering is left with the host,
which would have to report something ordered for anybody to have it.

This is issue 145 arriving by my own door an hour after I closed it: a
report that confidently says the opposite of the truth is worse than one
that says less.
2026-09-30 09:29:40 +02:00
jschoubben e10084fed6 Merge pull request 'Issue 087: a machine states its host only when the mesh sends it something' (#205) from issue/087-a-machine-says-its-host-only-when-asked into main 2026-09-30 07:09:17 +00:00
jschoubben 941b920bd7 Issue 087: a machine states its host only when the mesh sends it something
Measured live. All four machines run the identical host binary — same
digest, installed within eighteen seconds — and at first only one reported
a version, which read as a difference where there was none.

A machine publishes a report after an apply. The five-minute reconcile
publishes nothing, because it is the machine keeping itself as declared
rather than answering anything. So a current, idle machine never says, and
the mesh cannot tell that from a machine running something ancient.
Confirmed by pushing: not reported, then 04a27ca.

Enough for the purpose, not enough for the claim. For deciding whether a
new declaration field is safe it is sufficient — pushing is what the mesh
is about to do, and the answer arrives with the act. For knowing what the
mesh runs it is not, and "not reported" is worded as "nobody has asked
recently" for that reason.

Making a heartbeat carry it would close the gap and would change what a
heartbeat is — a bare word that the node is there, deliberately carrying
nothing else. Left alone rather than widened in passing.
2026-09-30 09:09:10 +02:00
jschoubben ecbd2ce2e6 Merge pull request 'Group 1: 145's report states its scope, and 107 waits for delivery' (#204) from issue/145-and-107-what-group-one-leaves into main 2026-09-30 07:00:32 +00:00
jschoubben b7f7b97d8a Group 1: 145's report states its scope, and 107 waits for delivery
145, partly resolved. The sentence that was true for eleven hours of a
mesh in which no module could reach another now says what it is not a
claim about: that is the mesh and the machines agreeing, and nothing here
dials a provision. It checks nothing and does not pretend to — ADR 0146
decides the check and is deliberately not built. What changed is that the
report no longer implies otherwise. Stays open for that reason.

Carried forward: 0146's check needs an internal name fetched over TLS with
the certificate verified, and until today no machine trusted the mesh's
authority. Three of four do now, so whoever builds it does not have to
solve that first.

107, diagnosed and deliberately not built. The premise is confirmed in the
host's own words — unknown fields are refused because "a field the host
does not know is a thing the control plane believes it asked for" — so the
fix is a flag day, not an addition. 087 now makes the cost measurable, and
the measurement is why it waits: one machine of four runs an older host,
it is parked, and nothing delivers a host at all (142). Shipping the field
means hand-placing binaries and unparking a machine, and one missed in
that sequence is unreachable, not degraded. The fault it prevents has
never been observed.

142 gains the note that it is 107's gate, and that it is what makes a
declaration field cost a rollout instead of an expedition.

A judgement about order, not a refusal, and cheap to overrule.
2026-09-30 09:00:18 +02:00
jschoubben c2cf0d72d6 Merge pull request 'Issue 087 is resolved: the mesh knows which host runs a machine' (#203) from issue/087-the-mesh-knows-which-host-runs-a-machine into main 2026-09-30 06:55:05 +00:00
jschoubben 0a5006b366 Issue 087 is resolved: the mesh knows which host runs a machine
The machine has reported its host version since ADR 0141, whose own
comment says why it must: without it nothing can say a machine is behind.
The controller's copy of the report did not have the field, so it
unmarshalled into nothing and was thrown away on arrival. Two structs
describe one message and only the sending side had it.

node show names it per machine, "not reported" where the mesh has not been
told. status names every machine running an older host than another does,
and which is newest.

Disagreement rather than staleness, deliberately: nothing delivers a host
version yet, so the mesh holds no canonical current one and "behind" has
no fixed point. What it can say is that the oldest host in the mesh is
what the mesh may send.

Two refusals to guess: a machine that reported nothing is not called
behind, and versions compare as strings — right for the timestamps this
mesh uses, wrong for a scheme where 10 sorts before 9, said at the place
that would have to learn.

107 gains the note that this is what makes its new field safe to consider,
and that one machine of four is behind today, so it is not free yet.
2026-09-30 08:54:52 +02:00
jschoubben 3ea9601904 Merge pull request 'Issue 125 is resolved: a hold is a line in the report, and it breaks "all well"' (#202) from issue/125-a-hold-is-a-line-in-the-report into main 2026-09-30 06:47:57 +00:00
jschoubben f7a37ee4f3 Issue 125 is resolved: a hold is a line in the report, and it breaks "all well"
Two of the four surfaces the report named already carried it — the host
has reported Held since ADR 0100, and node show reads the machine's own
list with an `as of` beside it. Recorded as checked rather than assumed.

Two did not. The apply line counted what it applied and said nothing
about the difference; status read the mesh's take-time listing, so a
module assigned after it showed nothing at all.

Both now say it, and the part that carries the weight: a hold suppresses
"all doing what they were told, all heard from, running what the mesh
would send them". That sentence was true for the whole outage, and acting
on it is what stopped the predecessor's proxy. Being adopted still does
not suppress it — a mode somebody chose is not a half-finished action.

Status does not call a hold a fault, deliberately. It is correct
behaviour, and a reader trained to see red for something the mesh did
right stops reading.
2026-09-30 08:47:36 +02:00
jschoubben 601d004fcb Merge pull request 'Issue 129: three machines of four trust the mesh, not one' (#201) from issue/129-three-machines-not-one into main 2026-09-30 00:30:30 +00:00
jschoubben d009c3efef Issue 129: three machines of four trust the mesh, not one
Extended after the first was proven. Every converged machine now holds
the anchor and verifies an internal name with a plain client; before,
the two unassigned ones answered 'unable to get local issuer
certificate' and held no entry for the mesh.

ace is excluded on purpose: it is adopted, so a module assigned there is
held rather than run, which is right and is not trust.

Both the resolution and 0147's insight said one machine of four, which
was true for about twenty minutes.
2026-09-30 02:30:23 +02:00
jschoubben 5036b927b9 Merge pull request 'Issue 129 is resolved: a workstation trusts the mesh, and stops when told to' (#200) from issue/129-a-machine-trusts-the-mesh into main 2026-09-30 00:18:51 +00:00
jschoubben 1f72e82b84 Issue 129 is resolved: a workstation trusts the mesh, and stops when told to
Registered ca-trust from the catalogue — it was merged and had never been
registered, which is why "assign it to one machine" had no module to name
— assigned it to the workstation, and verified.

Verified in the form ADR 0147 prescribes, against the authority's own API
so the handshake needs nothing else in the mesh to be right: 200, issuer
Mesh Internal CA, Verify return code 0. Four routed internal names verify
too, and `git ls-remote https://…` works, which is the consequence the
report named.

Removal exercised for the first time. Unassign and push removes the
anchor, empties the trust store of the mesh's authority, and returns the
plain client to the original error; assigning again restores it. That is
the half 0147 claimed and nothing had shown.

One thing it found that is not in the module: removal works only because
the host removes the service before the script. Stopping the unit is what
deletes the certificate and refreshes the bundles, and it needs the script
to still exist. The symmetry rests on an ordering nothing states.

0147's "written, and not yet run" now carries a progressive insight
saying it has run, where, and that it ran on one machine of four.
2026-09-30 02:18:21 +02:00
jschoubben 76fbe323ea Merge pull request 'Issue 118 is resolved: it was issue 135, and umami is healthy' (#199) from issue/118-is-135-and-is-resolved into main 2026-09-29 23:54:39 +00:00
jschoubben 295bf2f3ab Issue 118 is resolved: it was issue 135, and umami is healthy
Verified on the machine: umami has zero restarts, applies all 26
migrations, reports the store up to date, and umami.novox.be answers 200
where it had answered 502 since 2026-09-25.

It was the same fault as 135, filed three days earlier and diagnosed
without either record noticing the other. 135's container IS umami — it
is named in 135's own evidence table, holding novox.internal:10.42.0.1
after the overlay range moved. The dial appeared to succeed and the first
real query timed out because the name pointed at an address that no
longer existed; the store, the path and the credential were all fine.

118's own reasoning is kept as a warning, because it is careful and
wrong: it argued path-MTU and conntrack, and it named the move that would
have found it — compare umami against a container created that week —
and did not make it. A stale name presents as a network fault, which is
why ADR 0148 stops copying names into containers rather than detecting
when the copies go bad.

Also: 118's located-in named mesh-catalog modules/umami, which was never
at fault; it names mesh-host's comparison now. And 135 gains the pointer
to 118, which is the back-reference I have now missed three times.
2026-09-30 01:54:19 +02:00
jschoubben e1b74810a8 Merge pull request 'Issue 129 is live and reproduced, and needs three steps rather than one' (#198) from issue/129-and-what-reproducing-it-found into main 2026-09-29 23:30:48 +00:00
jschoubben e41eed0852 Issue 129 is live and reproduced, and needs three steps rather than one
The certificate is genuine, from Mesh Internal CA, and nothing on the
workstation trusts it — verbatim the error the report gives. The public
name on the same proxy verifies cleanly, which puts the fault exactly
where the report puts it.

What is in the way is not an assignment. `ca-trust` is merged in the
catalogue and has never been registered with the mesh — 39 of 76
manifests are — so there is no module to assign. It dry-runs clean and
needs no artifact built.

Two findings from reproducing it, both their own issues:

157 — every routed name is published with an `.internal` alias that
nothing serves. The hosts file says keycloak.novox.be.internal; the proxy
serves keycloak.novox.internal and refuses the other by name. The first
three names I tried came from the hosts file and failed with a TLS alert
rather than a verification error, which pointed at a regression that had
not happened.

158 — the proxy re-logs all 52 routes every two seconds, 31 times a
minute. The one line that explained 157 sat between two of them.

Also recorded, because it was nearly filed as a defect and is not one:
step-ca publishes roots as /roots.pem, which is PEM, so ca-trust's fetch
and its refuse-a-non-certificate guard are both right. Its other endpoint
/roots returns JSON that contains the text the guard greps for, so the
guard is sound only because of which path is published.
2026-09-30 01:30:24 +02:00
jschoubben ba30286896 Merge pull request 'The pointers back from what yesterday's records changed, which I missed twice' (#197) from decision/the-pointers-back-from-what-these-narrow into main 2026-09-29 22:47:45 +00:00
jschoubben 53b94c51bb The pointers back from what yesterday's records changed, which I missed twice
Three records were left describing a mechanism a new record had moved,
and a reader arrives at them by following a citation: 0066 still said a
routed name is written into every container after 0148 replaced that with
resolution; 0016 still read as though the lab were the test bed after
0149; and issues 109 and 135 said nothing about 0148 ending the copying
that 135's own fix made comparable. Each was a citation leading to the
wrong answer in a record that was not wrong about anything it decided.

This is the second time in one session. The playbook rule I added last
round did not stop it, so the convention is now written where the record
conventions live, with the shape to use and three worked examples — and
with the honest note that it is NOT machine-checked and cannot be from
`extends:` alone: 102 records extend another, 87 have no back-reference,
and that is correct, because extending usually means building on a
context. Making it mechanical means a record declaring the relationship
in frontmatter, which is a schema change and is not mine to decide.

Also: designs 18 and 20 claimed `updated:` dates from before I edited
them, and 117's `fixed-by` gained the commit beside the record.
2026-09-30 00:47:19 +02:00
jschoubben 83791f0921 Merge pull request 'The four open design questions, answered: ADRs 0148, 0149, 0150, and 0114 accepted' (#196) from decision/0148-the-meshs-names-are-resolved-not-copied into main 2026-09-29 22:39:38 +00:00
jschoubben bf4d4e2e7b The three parked questions are answered: 0114 accepted, 0068 superseded, 117 settled
**0114 accepted.** The separation it draws — a consumer's resource and a
credential that reaches it are different things — is a data-loss rule,
and a record naming one should not sit unresolved while the code that
could hit it is written. Not built, and accepting it schedules nothing;
accepted-and-not-built is where 0141 and 0142 already are.

**0068 superseded by 0149, the live mesh is the test bed.** Never built,
and contradicted by practice that was written down nowhere but a handoff
note. The faults that cost the most are faults of a mesh that already
exists — bound consumers, containers made against an older roster, an
adopted machine — and a bed is by construction a mesh that does not.
Issue 156 settles it: that change was verified honestly, on the only path
where it cannot fail. The lab is not retired; raising a mesh from bare is
now its whole job.

**117 answered by 0150.** A module's own code runs as supervised
processes under the module's one account. 0047's argument was for a
runtime per module, not for a container — a unit satisfies it and runs as
an account. The invariant is the account, not the process count: its
worry was a second identity to scope and seal, and processes sharing one
account create none. Designs 18 and 20 now cite it and 0047 carries a
dated pointer, which closes the third disagreement — that neither design
knew the record existed.
2026-09-30 00:39:09 +02:00
jschoubben ec42ee0846 ADR 0148: the mesh's names are resolved, not copied into every container
Answers issue 151. Copying the roster into each container made the roster
part of each container's identity, so one name moving replaced every
container in the mesh — and it never stopped the staleness it was for,
since a copy taken at creation is stale the moment the roster moves (109,
135).

A container resolves through its machine's resolver instead, and nothing
is copied. Staleness stops being possible rather than detected, and a
name's blast radius becomes nothing.

Scoping each container to the names it binds was the close call and is
rejected: it contradicts anything-calls-anything, and leaves the roster
in the digest so the churn returns for a widely-bound name.

Gated on issue 110 — a container on the runtime's default network has no
DNS at all today. Removing the copy first reintroduces 109 and 135
silently on a live mesh. 151 stays open until the code lands; design 08's
file-not-resolver passage is narrowed to the machine's own roster.
2026-09-30 00:34:30 +02:00
jschoubben 5a3dee9e9e Merge pull request 'The records pointed at branches that no longer exist, and two fixes had no sequel' (#195) from issue/pointers-that-resolve-to-nothing into main 2026-09-29 22:29:02 +00:00
jschoubben 47909c6b71 The records pointed at branches that no longer exist, and two fixes had no sequel
Three issues named the branch that fixed them, and a branch is deleted
when it merges — so every `fixed-by:` was a pointer that resolved to
nothing by the time anyone followed it. They name commits and pull
requests now, and playbook 03 says to.

Two records were missing the thing a reader arrives for. 146 did not say
that one of its fixes crash-looped the control plane on a running mesh,
which is the whole reason the delivery subject carries the stream and the
raise path was the only one exercised. 151 did not say that 152 removed
the false reasons its roster moved, or that it stays open for the real
ones.

ADR 0080 enumerates what cycle.py enforces and named four things; it
enforces five. A progressive insight names the fifth — the decision
stands, the list had gone stale. The checks README and playbook 03 gained
the same rule, and 155 points at all three.
2026-09-30 00:28:35 +02:00
jschoubben 7e5edab8da Merge pull request 'Issue 156: the wider grant has landed, and the notes the fix printed' (#194) from issue/156-the-grant-has-since-landed into main 2026-09-29 22:19:35 +00:00
jschoubben 3649f82204 Issue 156: the wider grant has landed, and the notes it printed
Measured on the broker's own config after the fixed controller composed
it: every node is now allowed _DELIVER.<node>.> as well as the bare
subject. That grant was missing when this was diagnosed, which is why
re-making the consumers would have silenced every machine.

Also records the five consumers it kept and reported, including the build
machine's worker, which is bound the same way and was not anticipated.
2026-09-30 00:18:57 +02:00
jschoubben b400ce2c54 Merge pull request 'Issue 156: moving a consumer's delivery subject stops a running mesh' (#193) from issue/156-a-consumer-that-works-is-not-replaced into main 2026-09-29 21:57:00 +00:00