ADR 0084 (extends 0027) — a provision is served by a node-scoped provider the
consumer selects, defaulting to co-location; a module may instead carry a private
embedded instance that is not a provision. Design: 01-to-be/23-choosing-a-provider.
ADR 0085 (extends 0031) — a secret is a provision and the vault is the module that
provides it; a module's own local secret becomes an ordinary pair credential that
rotates through the existing machinery, while the controller's provisioning-credential
mint (0048) is unchanged. Design: 01-to-be/24-the-secrets-vault; doc 13 amended to
cross-link the non-pair secret.
Issues 067/068 marked resolved with amended-design set. ADR index regenerated;
records and index checks pass.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Store and broker seats were made ordinary modules; secret-minting is still a
privileged property of the controller that no module owns. Propose the vault
become a module that provides a secret provision (generate/hold/rotate/backup/
audit), node-scoped like every other provider (issue 067), subsuming the three
secret paths and giving local-secret rotation a home.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Naming which provider is only half of how a module gets a database. The other
half: a module may carry its own version/fork-pinned instance, module-network
only, no published port, not a provision — and the model has no word for it.
Add it as a second axis with its own open questions.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The mesh models provisions as mesh-scoped (one provider of a kind, a single
mesh-store). But node-specific services delivered to the mesh was the plan from
the start: both nodes already run their own postgres, SQL server, redis and
object store, and identity — currently single — already serves apps on a second
node. The model cannot express which provider serves a consumer, so it collapses
a deliberately per-node fleet to one. Provider scoping is a whole-mesh decision
across postgres/s3-bucket/oidc, not an SSO patch.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Surfaced while reviewing the host apply loop for the controller/host
split discussion. 065: a permanently-failing resource retries for ever
with no escalation — reported, but never raises its hand as stuck.
066: a partly-applied declaration leaves a mixed state with no rollback,
which is harmless for independent resources and unexamined for pairs
that are only correct together. Both are questions HQ must answer, not
incidents — status open, no fix proposed.
The P1-sweep PR left both at 'located'; the fixes have since merged
(mesh-controller #31, mesh-tools #10) and the bed proves them. Flip to
resolved with their fixed-by.
060: 44 of 70 modules now mesh-buildable (was 8) — the mechanical
majority, proven by direct builds and the bed. The structural
remainders: external-dependency fetch (new issue 064) and route-proxy's
cross-repo build context. 064 records the build environment's isolation
from public npm and Docker Hub.
The broker's port was accepted on input but not forwarded, so a joined
node reached a DNAT'd broker only until it restarted. Fixed in
mesh-controller (25e42b3), proven by the built-store-cross-node bed
adopting the foundation's broker (which restarts it) and the joined
node still receiving declarations.
One push leaves the mesh consistent (the 057 decision, proposed for
acceptance); the shared runtime waits for its broker (058). Fixes on
mesh-control fix/one-push-is-enough and mesh-tools
fix/the-runtime-waits-for-its-broker; the built-store-cross-node bed
enforces both.
Fixed by mesh-controller PR 30 (79a9e17), rebased onto the merged
registry-reach train and proven by a green fresh run of the no-fake
two-node bed built from that commit.
One green fresh run of the no-fake two-node bed is the proof: node2's
consumers open the store and broker the mesh built and adopted, over
the overlay. Fixed by mesh-controller PR 29, mesh-catalog PR 26, gated
by mesh-lab PR 34.
061: the broker module's provisioner never ran; its runtime container
named no command and the image default is the tool host. 062: a failed
artifact-store lookup composed the network without the registry trust,
turning a transient error into permanent silent state. Both located,
fixes on the 042/048 train branches.
Only 8 of 70 modules carry a build section; the other 62 exist through an out-of-band
build script plus placeholder rewriting, so the mesh's own build-and-deliver path has
never run for them. Found by the no-fake bed; blocks it and the migration's delivery
assumption.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Findings of an adversarial review of the 055 fix: hub!=broker conflated, membership tested
as has-address, silent staleness, portless silent fallback, two hubs unrefused. One review
claim recorded as disputed against live wire measurements.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The glossary's authority page still named the controller's seat the-controller in two
entries, contradicting its own seat section after ADR 0079; issue 058's heading kept the
pre-renumber 059; 055's fixed-by named branches that stop existing after merge (now merge
commits/PRs) and its located-in listed file paths where the convention wants repos; 056's
located-in named mesh-host, which received no fix, instead of mesh-catalog; and the design
layer never said the one-store/one-broker property is enforced — 07-the-foundation and the
installation table now state the seats.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Observed in the two-node bed: a cross-node runtime exits on 'timed out fetching the
broker's certificate' until the tunnel forms, then settles. The retry belongs in-process.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The store and broker were "one per mesh" by convention only. Each foundation module now
claims a mesh-scoped seat named after the server it guards — postgres/mesh-store,
lavinmq/mesh-broker — and the controller's seat is renamed the-controller -> mesh-controller
so all three follow one rule. The resolver refuses a second holder, closing 056. Glossary,
the foundation and installation docs, and the decisions index follow the new name.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
055 is proven end-to-end by the two-node bed: a consumer on a joined node reaches the
adopted broker over the overlay, its binding names the control-node, and its vhost is
minted. Three fixes — the broker's amqps port in the firewall, the broker credential
naming the overlay not the public address, and pushing the provider node after the remote
consumer arrives. The last is an operator ordering, not a code fix, and its silent-failure
edge (a cross-node consumer that never provisions and only crash-loops) is opened as 057.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A two-node bed pins it: the firewall gap (broker's 5671 not in listens) is fixed
in mesh-catalog; the real bug is modules.go:294/build.go:231 building the broker
URL with the genesis public MESH_BROKER_ADDRESS, which a joined node cannot reach.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
"One postgres, one lavinmq" (ADR 0078) holds only by genesis assigning the
adopted modules to the control-node alone; nothing enforces it. A second assign
raises a divergent second server, silently. Open questions: a mesh-scoped
exclusive seat, or explicit adoption the controller refuses to place elsewhere.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Settles the design repository now that the self-upgrade build is on main:
- Records the two decisions that shipped without a record — ADR 0077 (the
controller/foundation/node vocabulary) and ADR 0078 (the store and broker are
ordinary modules); accepts ADR 0075 and 0076, which shipped work rests on.
- Fills issue 051's amended-design and wires ADR 0078 into 07-the-foundation.
- Sweeps the repo rename (mesh-control -> mesh-controller) into the mutable docs
now that the forge repo is renamed; updates the glossary note and repos.md.
- Fixes the six broken links from the design-doc renames, indexes the glossary,
regenerates the decisions reading order.
Both checks (records.py, index.py) are green. Statuses stay honest: the build is
on main and lab-proven but not deployed as the production mesh, so the to-be docs
remain in-progress and the as-is layer (the hal mesh) is unchanged — graduation
to implemented + as-is belongs to deployment, not merge.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Marks WBS 3.2/3.3/3.4 done and resolves issue 051: the foundation's store and
broker are adopted in place as the postgres and lavinmq modules, upgradeable
through their stated windows, source-tracked by status. A bare-metal mesh runs
one postgres and one lavinmq, proven 22/22 in the one-node lab. The two follow-up
gaps are tracked as issues 054 and 055.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
054: the adopted servers bind 0.0.0.0 from genesis but the packet filter is
installed later, so there is a window where they are open with only bootstrap
credentials. 055: the servers bind on the control-node and the one-node bed
cannot prove a consumer on another machine can reach them over the overlay.
Both are follow-ups to issue 051's adoption (WBS Phase 3), tracked rather than
rushed.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The lavinmq module raises its own server container and provisions vhosts on it,
separate from the substrate's mesh-broker. A mesh with the module assigned runs
two LavinMQ servers where one belongs — the exact AMQP twin of the two postgres
containers.
The module's own provisioner already assumes one server: it creates a vhost per
consumer, named for the login, isolated by the vhost boundary — the analog of
postgres's database-per-login. So the mesh's own control traffic is the / vhost
and every consumer's broker is a vhost beside it, all on one server.
Adoption therefore means the module does not run a server of its own: its server
is the substrate broker, adopted, and the module contributes the provisioner,
tools, events and bootstrap against it. The isolation model is already built;
what remains to design is only raising the one broker at genesis and then holding
it as a module, over the broker it is.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The substrate's store and the postgres module are the same thing. They were two
rows only while the substrate was a different KIND of thing — a store raised from
a bundle cannot provide postgres-database, so anything wanting a database needed
a second server. It is visible on any mesh built today: mesh-store and postgres,
two containers, the same image.
Which name survives is settled by the naming rule. Where a consumer speaks a
protocol the interface is the protocol, and "database" is not a capability. The
control plane's own queries use distinct on and on conflict, so the coupling is
to postgres and a "store" module would advertise a swap that fails the first time
anyone tries it.
The broker collapses the same way with a different outcome: amqp IS a protocol
that several implementations speak, so amqp is a legitimate provision and lavinmq
is one provider of it.
So adopting the substrate is not only an upgrade path — it is two rows of a
mesh's module list becoming one, twice. And it leaves nothing that is a specialty
after the pivot, which is the claim the whole design rests on and is not true
today for exactly those two.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
ADR 0014 is accepted and unambiguous: each module consumes its dependencies from
the private registry, the mesh's own shared library included, and a cross-package
change is publish then consume.
So the git dependency at a pinned commit is not a mechanism under consideration.
It is the shared library being consumed a way the record rules out, and the lock
naming a sibling directory is what that looks like when nobody publishes. Issue
053 is reclassified from a question about mechanism to a violation with a
direction.
What stays open is narrower and real: which software serves the private registry,
and that a fresh mesh has none when the first SDK is built — ADR 0014 assumes one
exists, and at genesis nothing has installed it.
Recorded after arguing at length for a bespoke content-addressed alternative,
against a decision that was already made and that I had not read.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The tool runtime's manifest names the SDK as a git dependency at a pinned commit;
its lock file names a sibling directory that exists on one workstation. It builds
only because the recipe runs npm install, which tolerates a lock disagreeing with
its manifest and re-resolves from the manifest — the one command that hides this.
It matters because a lock exists to make a build reproducible and this one
describes one machine, and because it is the first thing a fresh mesh builds: the
toolchain carries the SDK and everything with code of its own compiles inside it,
so a dependency resolved differently on the build machine than on a workstation
is a difference in every module the mesh will ever build.
And it is about to be copied. Each language's toolchain will carry that
language's SDK the same way, so the shape is worth settling before there are four
of them.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The packet filter generates its rules from what modules declare they listen on.
The broker is not a module, so it declares nothing, so its port is not opened.
Every machine dials that port to enrol and to receive every declaration it is
ever sent.
Invisible where it is assembled and fatal on the next machine: a mesh of one
never dials its own broker across the network, so the ruleset looks right. The
first machine to join is refused at the packet filter during enrolment, several
steps from anything that reports it — and assigning the firewall before joining
machines is both the natural order and the one that breaks.
This is 051 in a second place. That issue says the substrate cannot be updated
because the mesh holds no record of it; the same absence means the firewall
cannot know it exists. Ssh is already a floor for the same reason — a machine
nobody can reach is a machine nobody can repair — and the broker may be the same
class of fact.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The store and broker come from a bundle the installer writes once, with images
pinned in it, and nothing can change them afterwards: no build, no version to be
behind, no roll-out, and no way to report being out of date, because the mesh
holds no record of them as modules at all.
That is backwards. They are what everything else depends on, so their updates
matter most, and they are the only things with no mechanism to deliver one. A
mesh with a year-old broker reports itself entirely current.
The fix probably already exists: the control plane is carried, raised and then
adopted as an ordinary module pinned to what is running. Nothing in that pattern
is specific to the control plane. It would also remove a duplication visible on
any one-node mesh — the same postgres image running twice, because a store that
cannot be a module cannot provide a database to anything.
Recorded with the two hard parts stated rather than waved at: upgrading a store
the control plane is reading from, and upgrading a broker over the broker.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A mesh raised from bare metal built six modules and its catalogue reported three:
exactly those built after it began running. Missing were the shared base, the
store, and the catalogue itself.
The hole is never random. On a fresh mesh the modules built before the catalogue
are by necessity the ones it needed in order to exist, so the foundation is
always what is absent, on every mesh, at the moment the graph is first populated.
It breaks the question the catalogue is for: build edges hang off the base, so a
catalogue with no record of it answers 'what must be rebuilt' confidently and
wrongly. And nothing reports the gap, because a catalogue cannot know what it was
never told — this was found by comparing its answer against what the mesh had
just been watched doing.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A module's broker account is scoped to what it declares it emits and consumes. A
tool call needs a reply queue, which that scope does not cover and should not. So
the account is right, the request is reasonable, and no account exists that can
make it — asking a module its own question, from its own container, with its own
credential, is refused.
It matters because a module's tools are its operator-facing surface: the
catalogue serves the five questions it exists to answer and nothing can reach
them. It is also why a running module keeps being mistaken for a working one — a
test that cannot ask anything checks a container is up, and that substitution has
hidden two faults this week.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The registry says every machine pulls from it and opens its port to the mesh for
that reason. A machine that tries is refused by its own container runtime: the
registry serves plain HTTP and anything but loopback is treated as HTTPS.
It has never failed, and that is the finding. Every proof that a machine can
fetch a mesh-built artifact was a proof about the machine that built it, where
the reference was loopback. The bed that uses a routable address gets away with
it because the harness writes the runtime's configuration before the mesh exists.
Kept separate from 042, which they are easy to confuse: that one is about not
being allowed to pull, this one about not being able to whatever the credential
says. Fixing 042 alone leaves a machine with a valid account for a registry its
runtime will not talk to.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The report covered ports the firewall cannot close. The matching fault is ports
it does close and should not: rules come from what modules declare, a migrating
mesh knows about almost nothing, and anything listening on the host is dropped.
No module declares an ssh port and the generator has no allowance for one. The
session that loads the rules survives on conntrack until it drops, and then the
machine is reached from a rescue console.
Rehearsed on three lab machines before doing it on an anchor: loading the mesh's
rules refuses an ordinary host port and leaves a published container port
reachable, same prober, same second.
It is deliberate — no forward chain, because dropping there would stop every
container — but traffic to a published port never reaches the input chain, so
the firewall is silent about it. The anchor publishes 38 such ports today, all
filtered by the system being replaced, so the cutover would open every one of
them while reporting a firewall that is up.
Found while migrating the first module. Five approaches were tried and observed
to fail, including resolving the index and naming the platform at both ends; a
speculative fix was written and reverted rather than shipped, because it did not
make the mirror work.
One module uses this today, which is why it went unnoticed — and it is the shape
most of a migration wants, because the services being moved are third-party.
The mesh could already say which modules a base change invalidated, and could
not do anything about it: each recipe named one particular copy of the base by
fingerprint, and rebuilding produced the old one. Worse, the copy each named
existed only inside a throwaway lab, so those three modules could not be built
anywhere at all — and the line each replaced had the same fault.
Cloned alone, built by the mesh, every module moved onto it, and the base then
changed for real: all three went stale naming what moved. A comment-only change
correctly makes nothing stale, which found a second bug — staleness compared
commits where it should compare artifacts.
The runtime every module compiles against cannot be built by the mesh, so the
one rule that would catch it moving can never fire. And a container keeps the
values it was created with, so two good applies can leave it running on neither.
A build machine was refused the build queue, and this was raised as a gap in what
a manifest can express. It is not: `builder issue` creates exactly that account,
three lines from the code being read at the time.
Kept rather than deleted, for the one real thing in it — the wrong verb succeeds
and reports success, producing an account that authenticates and can do nothing,
so the failure surfaces a layer away as a permissions error that reads like a
missing feature.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Found by deleting the lab's registry and giving the machines a real path out:
public images fetched, the operator's own could not be fetched at all. There is
no provision for a registry credential, no manifest field, and no step in
enrolment that establishes one.
It applies to the mesh's own store too, which today asks for nothing — a
decision that has never been written down as one, and so reads as an absence.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Nine references now name the digest their tag resolved to. What closed is the
immediate fault; the open questions stand, because a digest in a repository is
wrong the moment anybody rebuilds — which is the reason the design wants the
repository to name artifacts and the mesh to hold digests.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Both carried a topic outside the six the index knows, so neither had a place to
be read in — 0067 had none at all. Both are 'the tiers', beside 0036 (bootstrap
ends at a usable mesh) and 0007 (connectivity), which is what they extend.
0067 cited 0041 for tier 0's property; on this trunk 0041 is events, and the
record it meant is 0005. A citation that resolves to the wrong record reads as
corroboration, which is worse than a dead link.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
There is no installer. The complete account of standing up a mesh is an
integration test in the lab, and a fixture may invent what it needs — this one
raised a registry no production has and rewrote every image reference through
it, concealing both 039 and the fact that a first node outside the lab had no
bootstrap path at all. The bed being green said nothing about whether a mesh
could be installed.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Nine images across seven modules name a tag, not a digest. ADR 0006 forbids it
and the host refuses it by name — and the refusal has never fired in a bed,
because the lab pushed every image into its own registry and rewrote every
reference to the digest it had just assigned. The harness was supplying the
property under test. Found by deleting the harness.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF