Design 24 flips to in-progress with mesh-catalog as its owner (playbook 04).
Starting the build surfaced two gaps the decision did not settle: a module
requiring `secret` receives exactly one value (069), and no command can
accept an operator's value into a consumer↔vault pair (070). Both opened as
issues rather than improvised around.
Also fills fixed-by on 067 and 068, which the cycle check refused as resolved
with no reference.
ADR 0084 (extends 0027) — a provision is served by a node-scoped provider the
consumer selects, defaulting to co-location; a module may instead carry a private
embedded instance that is not a provision. Design: 01-to-be/23-choosing-a-provider.
ADR 0085 (extends 0031) — a secret is a provision and the vault is the module that
provides it; a module's own local secret becomes an ordinary pair credential that
rotates through the existing machinery, while the controller's provisioning-credential
mint (0048) is unchanged. Design: 01-to-be/24-the-secrets-vault; doc 13 amended to
cross-link the non-pair secret.
Issues 067/068 marked resolved with amended-design set. ADR index regenerated;
records and index checks pass.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Store and broker seats were made ordinary modules; secret-minting is still a
privileged property of the controller that no module owns. Propose the vault
become a module that provides a secret provision (generate/hold/rotate/backup/
audit), node-scoped like every other provider (issue 067), subsuming the three
secret paths and giving local-secret rotation a home.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Naming which provider is only half of how a module gets a database. The other
half: a module may carry its own version/fork-pinned instance, module-network
only, no published port, not a provision — and the model has no word for it.
Add it as a second axis with its own open questions.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The mesh models provisions as mesh-scoped (one provider of a kind, a single
mesh-store). But node-specific services delivered to the mesh was the plan from
the start: both nodes already run their own postgres, SQL server, redis and
object store, and identity — currently single — already serves apps on a second
node. The model cannot express which provider serves a consumer, so it collapses
a deliberately per-node fleet to one. Provider scoping is a whole-mesh decision
across postgres/s3-bucket/oidc, not an SSO patch.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Surfaced while reviewing the host apply loop for the controller/host
split discussion. 065: a permanently-failing resource retries for ever
with no escalation — reported, but never raises its hand as stuck.
066: a partly-applied declaration leaves a mixed state with no rollback,
which is harmless for independent resources and unexamined for pairs
that are only correct together. Both are questions HQ must answer, not
incidents — status open, no fix proposed.
The P1-sweep PR left both at 'located'; the fixes have since merged
(mesh-controller #31, mesh-tools #10) and the bed proves them. Flip to
resolved with their fixed-by.
060: 44 of 70 modules now mesh-buildable (was 8) — the mechanical
majority, proven by direct builds and the bed. The structural
remainders: external-dependency fetch (new issue 064) and route-proxy's
cross-repo build context. 064 records the build environment's isolation
from public npm and Docker Hub.
The broker's port was accepted on input but not forwarded, so a joined
node reached a DNAT'd broker only until it restarted. Fixed in
mesh-controller (25e42b3), proven by the built-store-cross-node bed
adopting the foundation's broker (which restarts it) and the joined
node still receiving declarations.
One push leaves the mesh consistent (the 057 decision, proposed for
acceptance); the shared runtime waits for its broker (058). Fixes on
mesh-control fix/one-push-is-enough and mesh-tools
fix/the-runtime-waits-for-its-broker; the built-store-cross-node bed
enforces both.
Fixed by mesh-controller PR 30 (79a9e17), rebased onto the merged
registry-reach train and proven by a green fresh run of the no-fake
two-node bed built from that commit.
One green fresh run of the no-fake two-node bed is the proof: node2's
consumers open the store and broker the mesh built and adopted, over
the overlay. Fixed by mesh-controller PR 29, mesh-catalog PR 26, gated
by mesh-lab PR 34.
061: the broker module's provisioner never ran; its runtime container
named no command and the image default is the tool host. 062: a failed
artifact-store lookup composed the network without the registry trust,
turning a transient error into permanent silent state. Both located,
fixes on the 042/048 train branches.
Only 8 of 70 modules carry a build section; the other 62 exist through an out-of-band
build script plus placeholder rewriting, so the mesh's own build-and-deliver path has
never run for them. Found by the no-fake bed; blocks it and the migration's delivery
assumption.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Decisions were the one link the cycle checks skipped, and measuring found 19 of 70 records
orphaned — the credential flow and the module-runtime cluster among them, which is how a
stale premise about a settled decision survived in working memory. cycle.py now refuses an
accepted record nothing cites; the 19 got true homes (design frontmatter, the playbook that
implements 0021, META for the process records). The overview names the practice: spec-driven
development with provenance.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Findings of an adversarial review of the 055 fix: hub!=broker conflated, membership tested
as has-address, silent staleness, portless silent fallback, two hubs unrefused. One review
claim recorded as disputed against live wire measurements.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The flow the process overview draws — idea/symptom -> decision -> to-be design -> code ->
as-is — was enforced by nothing. cycle.py now refuses a to-be design naming no decision, an
in-progress/implemented design naming no owning code, a located/fixed issue with no owner,
a fixed/resolved issue with no fix, and a graduated research overview that does not say what
it became. AGENTS.md carries the cycle and a where-to-look table so a fresh session (or a
cleared context) finds the chain in frontmatter instead of assuming it. Grounding the check
surfaced two real gaps, fixed here: the work-ahead design named no owning code, and research
003 listed one became target twice.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The glossary's authority page still named the controller's seat the-controller in two
entries, contradicting its own seat section after ADR 0079; issue 058's heading kept the
pre-renumber 059; 055's fixed-by named branches that stop existing after merge (now merge
commits/PRs) and its located-in listed file paths where the convention wants repos; 056's
located-in named mesh-host, which received no fix, instead of mesh-catalog; and the design
layer never said the one-store/one-broker property is enforced — 07-the-foundation and the
installation table now state the seats.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Observed in the two-node bed: a cross-node runtime exits on 'timed out fetching the
broker's certificate' until the tunnel forms, then settles. The retry belongs in-process.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The store and broker were "one per mesh" by convention only. Each foundation module now
claims a mesh-scoped seat named after the server it guards — postgres/mesh-store,
lavinmq/mesh-broker — and the controller's seat is renamed the-controller -> mesh-controller
so all three follow one rule. The resolver refuses a second holder, closing 056. Glossary,
the foundation and installation docs, and the decisions index follow the new name.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
055 is proven end-to-end by the two-node bed: a consumer on a joined node reaches the
adopted broker over the overlay, its binding names the control-node, and its vhost is
minted. Three fixes — the broker's amqps port in the firewall, the broker credential
naming the overlay not the public address, and pushing the provider node after the remote
consumer arrives. The last is an operator ordering, not a code fix, and its silent-failure
edge (a cross-node consumer that never provisions and only crash-loops) is opened as 057.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A two-node bed pins it: the firewall gap (broker's 5671 not in listens) is fixed
in mesh-catalog; the real bug is modules.go:294/build.go:231 building the broker
URL with the genesis public MESH_BROKER_ADDRESS, which a joined node cannot reach.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
"One postgres, one lavinmq" (ADR 0078) holds only by genesis assigning the
adopted modules to the control-node alone; nothing enforces it. A second assign
raises a divergent second server, silently. Open questions: a mesh-scoped
exclusive seat, or explicit adoption the controller refuses to place elsewhere.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Settles the design repository now that the self-upgrade build is on main:
- Records the two decisions that shipped without a record — ADR 0077 (the
controller/foundation/node vocabulary) and ADR 0078 (the store and broker are
ordinary modules); accepts ADR 0075 and 0076, which shipped work rests on.
- Fills issue 051's amended-design and wires ADR 0078 into 07-the-foundation.
- Sweeps the repo rename (mesh-control -> mesh-controller) into the mutable docs
now that the forge repo is renamed; updates the glossary note and repos.md.
- Fixes the six broken links from the design-doc renames, indexes the glossary,
regenerates the decisions reading order.
Both checks (records.py, index.py) are green. Statuses stay honest: the build is
on main and lab-proven but not deployed as the production mesh, so the to-be docs
remain in-progress and the as-is layer (the hal mesh) is unchanged — graduation
to implemented + as-is belongs to deployment, not merge.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Marks WBS 3.2/3.3/3.4 done and resolves issue 051: the foundation's store and
broker are adopted in place as the postgres and lavinmq modules, upgradeable
through their stated windows, source-tracked by status. A bare-metal mesh runs
one postgres and one lavinmq, proven 22/22 in the one-node lab. The two follow-up
gaps are tracked as issues 054 and 055.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
054: the adopted servers bind 0.0.0.0 from genesis but the packet filter is
installed later, so there is a window where they are open with only bootstrap
credentials. 055: the servers bind on the control-node and the one-node bed
cannot prove a consumer on another machine can reach them over the overlay.
Both are follow-ups to issue 051's adoption (WBS Phase 3), tracked rather than
rushed.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
"control plane" -> controller and "substrate" -> foundation throughout
03-DESIGN, 00-META and the README, with 06-the-control-plane.md and
07-the-substrate.md renamed to 06-the-controller.md and 07-the-foundation.md.
The immutable 02-DECISIONS records keep their original wording (and links to
them are unchanged) — a term retired here may still appear there, which the
glossary explains how to read.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Locks the vocabulary that kept drifting in conversation — controller (not
"control plane"), foundation (not "substrate"), node and control-node, seat /
bench / claim, package vs artifact. AGENTS.md points at it as the authority.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Records the decision the package-registry work turns on — the SDK is built on a
public base and published before the toolchain that consumes it, so nothing is
circular; mesh-tools stays the thin toolchain base but resolves the SDK by
version. Reconciles docs 12/17/22 and indexes the record.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The protocol spec claimed the envelope and grant drifted across implementations.
Inspection showed the wire agrees — envelope required headers match, the two
optional ones are legitimately optional, and the grant wire (the contributions
file) is identical on both sides. The disagreement was in dead types, now removed.
So phase 1's 'make them agree' work is done by deletion and correction. What
remains is a conformance fixture as prevention — pinning the envelope and the
contributions file so a future change that breaks agreement fails a test — and a
full per-capability suite is deferred until a third language actually needs it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The record claimed the two implementations already disagreed. Inspection showed
the live wire agrees: the disagreeing grant types were dead (removed), and the
envelope's two extra headers are optional and set when relevant, not missing.
The danger was dead types contradicting the wire, not live disagreement — which
is a sharper reason for specifying the wire and checking against it, not a weaker
one. The model stands; the conformance suite's job is prevention rather than
repairing a present break.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The correction the operator pushed: stop running a 20-minute lab against a mesh
mid-transformation, debugging paths the next phase deletes. The base-build hang
is almost certainly the SDK resolving from a git URL inside a docker build (issue
053), which Phase 2 removes — so debugging it on the current shape is debugging
deprecated code.
Phase 0 folds in: the installer's own regressions are fixed and committed;
whether it runs green is the final acceptance test, after the phases that change
its build path are in. Faults that can be reasoned out of the code path are, by
reading rather than running.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Everything decided this cycle and not yet built. Phase 0 gets the installer
green, because nothing else is testable end to end without it. Phase 1 makes the
protocol one thing and fixes the Go/TS drift the installer's own provisioning
exercises. Phase 2 stands up the private package registry ADR 0014 assumes and
publishes the SDK into it. Phase 3 adopts the substrate so one postgres and one
lavinmq serve everything, which is the hardest and needs all three above.
Order is dependency, not preference. Each phase ends at a run rather than a
paragraph, because a phase that ends at a claim is how things went missing this
cycle without anything complaining.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The lavinmq module raises its own server container and provisions vhosts on it,
separate from the substrate's mesh-broker. A mesh with the module assigned runs
two LavinMQ servers where one belongs — the exact AMQP twin of the two postgres
containers.
The module's own provisioner already assumes one server: it creates a vhost per
consumer, named for the login, isolated by the vhost boundary — the analog of
postgres's database-per-login. So the mesh's own control traffic is the / vhost
and every consumer's broker is a vhost beside it, all on one server.
Adoption therefore means the module does not run a server of its own: its server
is the substrate broker, adopted, and the module contributes the provisioner,
tools, events and bootstrap against it. The isolation model is already built;
what remains to design is only raising the one broker at genesis and then holding
it as a module, over the broker it is.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The substrate's store and the postgres module are the same thing. They were two
rows only while the substrate was a different KIND of thing — a store raised from
a bundle cannot provide postgres-database, so anything wanting a database needed
a second server. It is visible on any mesh built today: mesh-store and postgres,
two containers, the same image.
Which name survives is settled by the naming rule. Where a consumer speaks a
protocol the interface is the protocol, and "database" is not a capability. The
control plane's own queries use distinct on and on conflict, so the coupling is
to postgres and a "store" module would advertise a swap that fails the first time
anyone tries it.
The broker collapses the same way with a different outcome: amqp IS a protocol
that several implementations speak, so amqp is a legitimate provision and lavinmq
is one provider of it.
So adopting the substrate is not only an upgrade path — it is two rows of a
mesh's module list becoming one, twice. And it leaves nothing that is a specialty
after the pivot, which is the claim the whole design rests on and is not true
today for exactly those two.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Reported as 'the host does not survive a reboot', which reads as a mesh that
cannot come back. mesh-host/packaging/ ships nox-mesh-host.service and two
companions. The installer declines to place them because a unit file is a
packaging decision, and the lab starts the host with --host-in-background, which
says in its own help that it does not survive a reboot.
So the lab run failing this was the lab being honest, and the gap is the step
that puts a shipped unit on a machine — narrower and more fixable than what I
wrote.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The first draft implied gitea and the registry could not share a machine, because
registry claims the-artifact-store at node scope and I carried that across to
gitea without asking what the claim is for.
A machine running gitea for git and packages alongside a registry serving
artifacts is an ordinary arrangement. They are different ports doing different
jobs, and nothing about one being the mesh's artifact store requires the other
not to exist.
The exclusivity that matters is mesh-wide and already expressed: provides at mesh
scope means two providers are two answers, and the resolver refuses until one is
assigned. Forbidding co-residence adds nothing and forbids something reasonable.
Whether registry should still hold that claim is left open rather than answered
from outside its manifest — it may be protecting something about its port or its
data directory that nobody wrote down.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Two questions circling turned out to be one asked twice: should gitea be the
mesh's registry, and where does the SDK come from. The framing that dissolves
both is that artifact-store is already a provision and the registry already
provides it — so this was never about replacing a component. It is a second
provider of an existing provision, which this mesh has a mechanism for and uses
for certificate authorities already.
So: two provisions, because they are two jobs. artifact-store is content
addressed, pinned by digest, no versions and no ranges — what the mesh delivers
to machines. package-registry is an ecosystem's own, addressed by name and
version — what code resolves when compiled. Conflating them is how a mesh that
pins everything ends up rebuilding one commit into two different things.
The small registry stays the provider genesis installs, not because it is better
but because of what it is: a directory and one container, installable where there
is no database and no control plane. Gitea needs both, and the pivot needs
somewhere to publish before either exists.
Gitea also provides artifact-store, so a mesh may choose it — and choosing it
answers issues 042 and 048 by adopting something that already has accounts and
TLS, rather than reimplementing them in a registry that has neither.
Moving between providers is a designed act with a verification step that is easy
to skip and is the only thing between it and a mesh that cannot restart its own
control plane.
And the bootstrap still has no package registry when the first build needs one.
Named rather than solved, so the next person does not discover it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Every step from a bare machine to a mesh that maintains itself, in three phases,
with each step named as the installer prints it.
The point of writing it out is the shape it exposes. The installer owns twelve
steps and ends at a mesh that RUNS. Seven more turn that into a mesh that WORKS —
the shared base, a store that is a provider rather than the control plane's own
memory, the catalogue, the replay of what was built before the catalogue existed,
the control plane rebuilt through the module path, the private network with the
node actually placed on it, and the packet filter. None of those seven is the
installer's. They are things somebody types, which is why a test had to be
written to discover they were missing.
Machines arrive last, in phase three, because a machine joining a mesh that
cannot build anything proves enrolment works and nothing else.
And five things that are not yet true are named rather than implied: phase two is
manual, a second machine cannot pull what the mesh built, ADR 0014 assumes a
private package registry that genesis has not installed when the first build
needs it, the host agent does not survive a reboot, and nothing can contradict a
claim that a machine was installed this way.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
ADR 0014 is accepted and unambiguous: each module consumes its dependencies from
the private registry, the mesh's own shared library included, and a cross-package
change is publish then consume.
So the git dependency at a pinned commit is not a mechanism under consideration.
It is the shared library being consumed a way the record rules out, and the lock
naming a sibling directory is what that looks like when nobody publishes. Issue
053 is reclassified from a question about mechanism to a violation with a
direction.
What stays open is narrower and real: which software serves the private registry,
and that a fresh mesh has none when the first SDK is built — ADR 0014 assumes one
exists, and at genesis nothing has installed it.
Recorded after arguing at length for a bespoke content-addressed alternative,
against a decision that was already made and that I had not read.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
What the SDK contains, answered by exclusion as much as by inclusion. It is the
protocol and nothing else — no configuration loader, because configuration
arrives as files the mesh wrote; no API clients, because a Plex client changes
when Plex changes and that has nothing to do with any other module; no storage,
HTTP or logging, because the language has those. The test for anything proposed
is ADR 0039's: does editing it recompile unrelated modules, and does it change
often. Both, and it stays out.
Then the worked module: events in TypeScript, tools in Go, a provisioner in Rust,
a scheduled job in Python. Four artifacts, four toolchains, four processes, one
module — and each part is an ordinary project in its language depending on the
mesh SDK the ordinary way, so a laptop resolves what a build resolves.
And publishing a package as a module capability, which makes the SDK unspecial:
it is simply the first module that published a library. A Plex client belongs to
the Plex module because that is the only thing that knows when Plex changed.
Three things left open rather than papered over: which registry (the catalogue
holds verdaccio and a forge usually serves one too, and nothing says which is
ours), who may publish (a credential that does not exist), and what a range means
in a mesh where everything else is pinned by digest — a mesh that can rebuild a
commit and get a different library is a real change, and should be decided rather
than arrived at.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A floor every implementation needs and three capabilities independent of each
other, so an SDK can implement the floor and events and be a real thing rather
than an unfinished one.
Written as a specification, which means it says what is required rather than how
anything is arranged — and says plainly where it describes behaviour that is not
yet true. Three places it does:
x-causation-id and x-schema are specified and emitted by nothing; the Go side
writes four headers and the TypeScript side declares six. A module may serve
tools and may not call them, because a caller needs a reply queue its account may
not declare. And the two implementations disagree about what a grant carries — in
TypeScript consumer is the module, in Go it is the node and the module is From.
One word, two meanings, in two halves of one mesh.
Naming those in the specification rather than leaving them for conformance to
discover, because a specification that only described what already works would
have nothing to say about the things most likely to break.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The tool runtime's manifest names the SDK as a git dependency at a pinned commit;
its lock file names a sibling directory that exists on one workstation. It builds
only because the recipe runs npm install, which tolerates a lock disagreeing with
its manifest and re-resolves from the manifest — the one command that hides this.
It matters because a lock exists to make a build reproducible and this one
describes one machine, and because it is the first thing a fresh mesh builds: the
toolchain carries the SDK and everything with code of its own compiles inside it,
so a dependency resolved differently on the build machine than on a workstation
is a difference in every module the mesh will ever build.
And it is about to be copied. Each language's toolchain will carry that
language's SDK the same way, so the shape is worth settling before there are four
of them.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Two corrections, both from the operator and both better than what was written.
An SDK is an implementation of the mesh's module protocol in one language, and
nothing more. The first draft defined it by the test it passes, which describes
how you check one rather than what one is — and leaves it sounding like a library
that helpers could accumulate in.
And the protocol is split per capability, which was missing entirely. A module
that only consumes events uses the event capability; one that serves tools uses
the tool capability; a provider uses provisioning. Nothing about consuming an
event requires knowing how a grant is answered, so an SDK need not implement all
of it to be real.
That has a precedent here: a host declares which resource kinds it can apply, and
a partial host is a real thing rather than a broken one (ADR 0005). An SDK
implementing the floor and events is exactly as legitimate, and a module written
against it is a module that does events.
Which changes what adding a language costs. A Rust SDK doing connection and
events is useful the day it exists, with tools and provisioning following when
something needs them — rather than a language being unsupported until it is
entirely supported, which is what makes adding one a project instead of a
contribution.
Conformance is therefore per capability too: a monolithic pass or fail would make
a partial implementation indistinguishable from a broken one, which is the
distinction the whole thing rests on.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
ADR 0039 settles what belongs in an SDK. It does not say what happens when there
is more than one, and there already is: the contracts are expressed as Go types
in the control plane and host and as TypeScript types in the SDK, and nobody has
felt it because both live in one repository.
They already disagree. The provision's field is "resource" in one and "Provision"
in the other; "consumer" means the module in one and the node in the other; the
envelope declares six headers on one side and emits four on both — the missing
two being x-causation-id and x-schema, the second of which is exactly what a body
needs in order to change shape without silent misreads.
That class of failure does not announce itself. Two implementations disagreeing
about an envelope do not fail to compile — they ignore each other's messages, and
a mesh where a module stops reacting looks like a mesh where nothing happened.
So the decision is to specify the wire rather than share the types, because the
shapes are the easy half. What two implementations actually disagree about is
behaviour: queue naming and durability, which headers are required and what an
unknown one means, taking identity from the sealed credential rather than the
environment, dedup on an id only the emitter can make, pinning a fingerprint
rather than trusting an authority.
And the suite is executable rather than prose, because a specification nobody can
run is a document two implementations drift from while both believe they conform.
The two existing implementations are the first made to pass it — a suite only new
SDKs must satisfy would certify every future language against a disagreement that
is already here.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A reference table goes stale the day somebody adds a field, so this one points at
mesh-catalog/modules/showcase — a module that uses all of it — and a test that
fails when it stops doing so. Read the module when the table disagrees with it.
Two rows in the coverage survey were stale because of this week's work: systemd
units were a file plus a service, which made every author write unit syntax and
is why "process" exists; and building from source was images only, where a
bundle now names a language and lets the mesh choose the toolchain.
And the two rows at the bottom of the resource table are the interesting ones.
"action" is refused to modules outright — the link may not carry a command, so a
module needing something done ships a program that reconciles. "service"
installs no unit by design, right for software shipping one and wrong for code
the mesh built, which has none until the mesh writes it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The first draft said the builder refused archives. It does not. An archive is
packed deterministically, hashed, published by digest, fetched by the machine and
unpacked — the whole path exists. Only the local builder used at genesis refuses
one, and deliberately: an archive is bytes that mean nothing until something
serves them, and at genesis nothing does.
What an archive cannot do is compile. Its source is a directory packed as it
stands, so shipping compiled output means compiling somewhere first, which means
a Dockerfile — the burden this document is about. The gap is not the artifact
kind. It is that no recipe both builds and packs.
Found by reading the builder rather than the manifest schema, which is where the
first draft's claim came from.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
12-a-module-repository says what a module may build and where it goes. Nothing
said how a build is MODELLED, and the model is the problem: a recipe is implicit,
singular and always a Dockerfile; a toolchain is not modelled at all, arriving as
two build arguments the module hand-writes; a language is not a concept; and an
archive is declared in the manifest and refused by the builder.
The cost is measurable rather than theoretical. Adding a module with its own code
means repeating an incantation - two ARG bases, a specific working directory so
the SDK resolves upward, the compiler invoked by absolute path because the usual
symlink is resolved away when the base is assembled, a second stage, an env var
naming the entrypoints. Most of the catalogue is unconverted, and two conversions
done in one session were each wrong twice with a working example open.
So: recipe becomes explicit with three kinds, and toolchain becomes derived from
a declared language rather than written by every author. A Dockerfile stays, and
stops being compulsory - it is right for software needing a particular base and
wrong for "compile my module's code", which is the same operation every time.
The cost is stated before it is chosen: every language is permanent, and the
contracts are already expressed twice - Go structs and TypeScript types kept in
step by hand. A second language makes that drift. So language-neutral contracts
come first, or the drift gets worse while hiding.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The packet filter generates its rules from what modules declare they listen on.
The broker is not a module, so it declares nothing, so its port is not opened.
Every machine dials that port to enrol and to receive every declaration it is
ever sent.
Invisible where it is assembled and fatal on the next machine: a mesh of one
never dials its own broker across the network, so the ruleset looks right. The
first machine to join is refused at the packet filter during enrolment, several
steps from anything that reports it — and assigning the firewall before joining
machines is both the natural order and the one that breaks.
This is 051 in a second place. That issue says the substrate cannot be updated
because the mesh holds no record of it; the same absence means the firewall
cannot know it exists. Ssh is already a floor for the same reason — a machine
nobody can reach is a machine nobody can repair — and the broker may be the same
class of fact.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The store and broker come from a bundle the installer writes once, with images
pinned in it, and nothing can change them afterwards: no build, no version to be
behind, no roll-out, and no way to report being out of date, because the mesh
holds no record of them as modules at all.
That is backwards. They are what everything else depends on, so their updates
matter most, and they are the only things with no mechanism to deliver one. A
mesh with a year-old broker reports itself entirely current.
The fix probably already exists: the control plane is carried, raised and then
adopted as an ordinary module pinned to what is running. Nothing in that pattern
is specific to the control plane. It would also remove a duplication visible on
any one-node mesh — the same postgres image running twice, because a store that
cannot be a module cannot provide a database to anything.
Recorded with the two hard parts stated rather than waved at: upgrading a store
the control plane is reading from, and upgrading a broker over the broker.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A mesh raised from bare metal built six modules and its catalogue reported three:
exactly those built after it began running. Missing were the shared base, the
store, and the catalogue itself.
The hole is never random. On a fresh mesh the modules built before the catalogue
are by necessity the ones it needed in order to exist, so the foundation is
always what is absent, on every mesh, at the moment the graph is first populated.
It breaks the question the catalogue is for: build edges hang off the base, so a
catalogue with no record of it answers 'what must be rebuilt' confidently and
wrongly. And nothing reports the gap, because a catalogue cannot know what it was
never told — this was found by comparing its answer against what the mesh had
just been watched doing.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A module's broker account is scoped to what it declares it emits and consumes. A
tool call needs a reply queue, which that scope does not cover and should not. So
the account is right, the request is reasonable, and no account exists that can
make it — asking a module its own question, from its own container, with its own
credential, is refused.
It matters because a module's tools are its operator-facing surface: the
catalogue serves the five questions it exists to answer and nothing can reach
them. It is also why a running module keeps being mistaken for a working one — a
test that cannot ask anything checks a container is up, and that substitution has
hidden two faults this week.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The registry says every machine pulls from it and opens its port to the mesh for
that reason. A machine that tries is refused by its own container runtime: the
registry serves plain HTTP and anything but loopback is treated as HTTPS.
It has never failed, and that is the finding. Every proof that a machine can
fetch a mesh-built artifact was a proof about the machine that built it, where
the reference was loopback. The bed that uses a routable address gets away with
it because the harness writes the runtime's configuration before the mesh exists.
Kept separate from 042, which they are easy to confuse: that one is about not
being allowed to pull, this one about not being able to whatever the credential
says. Fixing 042 alone leaves a machine with a valid account for a registry its
runtime will not talk to.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The report covered ports the firewall cannot close. The matching fault is ports
it does close and should not: rules come from what modules declare, a migrating
mesh knows about almost nothing, and anything listening on the host is dropped.
No module declares an ssh port and the generator has no allowance for one. The
session that loads the rules survives on conntrack until it drops, and then the
machine is reached from a rescue console.
Rehearsed on three lab machines before doing it on an anchor: loading the mesh's
rules refuses an ordinary host port and leaves a published container port
reachable, same prober, same second.
It is deliberate — no forward chain, because dropping there would stop every
container — but traffic to a published port never reaches the input chain, so
the firewall is silent about it. The anchor publishes 38 such ports today, all
filtered by the system being replaced, so the cutover would open every one of
them while reporting a firewall that is up.
Found while migrating the first module. Five approaches were tried and observed
to fail, including resolving the index and naming the platform at both ends; a
speculative fix was written and reverted rather than shipped, because it did not
make the mirror work.
One module uses this today, which is why it went unnoticed — and it is the shape
most of a migration wants, because the services being moved are third-party.
The design said how the builder arrives was unsettled and that nothing
installed it — the one gap stopping a fresh mesh from producing anything. Both
are now false. What is still true is narrower and worth keeping separate:
nothing asks a raised mesh for the rest of the catalogue, and no bed asserts
that it could.
The mesh could already say which modules a base change invalidated, and could
not do anything about it: each recipe named one particular copy of the base by
fingerprint, and rebuilding produced the old one. Worse, the copy each named
existed only inside a throwaway lab, so those three modules could not be built
anywhere at all — and the line each replaced had the same fault.
The builder's arrival was the one rule the document said nothing checked. It is
checked now, by both genesis beds — and so is the thing that distinguishes a
built control plane from a carried one, which every earlier assertion accepted
either way.
The section saying the change was decided and had not happened now contradicted
the section below it. It also records the argument that failed, because a reader
will otherwise ask the same question and reach the same wrong answer.
Written an hour ago claiming a produced image must be published before anything
can fetch it, so the registry had to precede the control plane. The premise is
false: the machine that builds the image is the machine that runs it, and the
temporary control plane names a built image exactly as it names a carried one —
by the digest of its own configuration, which requires nothing to have served
it. Building changes where the bytes came from, not where they are.
Rewritten rather than superseded because nothing has been built on it and
nobody has read it: a record that contradicts itself is a draft, not a decision.
The argument is kept, because it was asked for and a negative answer is the
result.
The two questions the design record named as the one gap stopping a fresh mesh
from producing anything. They cannot be answered apart: a builder with nowhere
to publish has made a file on a disk.
The registry's role did not change — the answer to 'must it precede the control
plane' did, because the control plane's image is now produced rather than
carried, and a produced image must be put somewhere before it can be fetched.
Cloned alone, built by the mesh, every module moved onto it, and the base then
changed for real: all three went stale naming what moved. A comment-only change
correctly makes nothing stale, which found a second bug — staleness compared
commits where it should compare artifacts.
The runtime every module compiles against cannot be built by the mesh, so the
one rule that would catch it moving can never fire. And a container keeps the
values it was created with, so two good applies can leave it running on neither.
Corrects ADR 0070, written an hour earlier, which had the control plane consuming
the catalogue in order to compose a declaration. That was written before the two
graphs had been told apart and creates a dependency that need not exist: a
catalogue that is down would leave the control plane unable to compose the thing
that would repair it.
The catalogue links module-versions to each other and does not know nodes exist.
The control plane links module-versions to nodes, and holds capabilities and
claims. They meet only when something is installed, and everything between them
travels as events over the broker, on durable queues, so nothing is lost when a
receiver is away.
Build order is not computed anywhere. The builder never consults the graph and
builds what it is asked for; the catalogue asks for the next build after the
previous registration, so ordering holds by construction. Its rule is a condition
rather than a schedule — rebuild once everything a module was built against is
current — which covers a chain and a diamond alike.
Left open: whether an upgrade is applied or merely noticed, and whether a module
on several machines upgrades on all at once.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
ADR 0070 has the init builder clone the source and does not say from where, and
ADR 0067 had rejected building at genesis partly because the forge runs on the
mesh being rebuilt. That objection binds only when those are the same mesh, which
is true exactly once.
So genesis clones from a mesh by name, and if that mesh is lost the name moves to
another that holds a copy — recovery is a name pointing elsewhere rather than a
backup being restored, and every installation adds somewhere it could point.
It names a commit and checks what it got, because the forge a mesh installs from
is the trust anchor for everything that mesh will run. On 2026-09-11 that forge
was running a cryptominer and tampering with git operations in flight; nothing was
altered, but a mesh installing during those hours could not have known that.
What relationship a mesh keeps afterwards is left open and named, so that whoever
writes the init builder does not settle it by accident: a snapshot and then
independence, or a continuing upstream for core modules.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx